PaperScope
LIVE · 2026-10-07 05:40 UTC

REViT-v2: Hierarchical Windowed Roto-reflection Equivariant ViT for Equivariant Feature Extraction

Sheir A. Zaheer, Jihwan Moon, Chan Y. Park

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.07585 v1
Category
Submitted
2026-10-06

Abstract

We propose a scalable roto-reflection-group-equivariant vision transformer based on windowed group-convolutional self-attention and a hierarchical feature architecture. We demonstrate that our approach can be scaled to group-equivariant vision transformers (ViTs) with millions of parameters and large datasets with practically sized images, i.e., ImageNet. The code and pretrained weights for the proposed Hierarchical Windowed Roto-reflection Equivariant ViTs (REViT-v2) are available at https://github.com/kc-ml2/revit.

Comment: 7 pages, Accepted for presentation at NeurIPS NeurREPS workshop 2026

arXiv abs page · PDF · same-day batch