VGGT-R: You Can Really Accelerate
Vision Geometry Grounded Transformer
with the Underestimated Registers

KargoBot  ·  University of Chinese Academy of Sciences  ·  Zerith Robotics
Tsinghua University  ·  Nanyang Technological University

* Equal contribution    † Corresponding author

Interactive 360-degree reconstruction of the 7Scenes Chess scene Drag to rotate
Interactive 360-degree reconstruction of the 7Scenes Red Kitchen scene Drag to rotate

Abstract

Corresponding input image sequence and VGGT-R geometric results

VGGT-R repurposes pretrained register tokens as compact cross-frame proxies, accelerating long-sequence 3D perception without retraining or explicit token pruning.

Large-scale 3D vision foundation models such as Visual Geometry Grounded Transformer (VGGT) deliver strong performance across geometric perception tasks, but the quadratic complexity of global attention severely limits long-sequence scalability. We present VGGT-R, a training-free framework that repurposes register tokens to reduce global attention overhead while preserving geometric performance. A fine-grained layer-wise analysis reveals that registers act as compact cross-frame proxies. We split the network into a geometric front-end with restricted attention and a geometric back-end with register-guided token selection. VGGT-R achieves 3.2× speedup without explicitly pruning tokens, while improving depth estimation by approximately 10% and remaining competitive across other geometric tasks.

Method

A layer-aware design that gives registers a different job at each stage of geometric reasoning.

Overall VGGT-R framework with restricted attention and register-guided token selection
Overall framework of VGGT-R. The geometric front-end routes cross-frame interaction through registers, while the back-end preserves full queries and selects informative keys and values.
01

Restricted Attention

Cross-frame communication is routed solely through camera and register tokens in selected front-end blocks, eliminating dense patch-to-patch interactions.

02

Register-Guided Selection

Three complementary signals preserve relevant, structurally distinctive and cross-frame novel patches in the back-end keys and values.

03

Training-Free Integration

The original token layout and prediction heads remain unchanged. No additional data, fine-tuning or explicit patch-token pruning is required.

Results

VGGT-R improves the accuracy–efficiency trade-off across depth, point clouds and camera pose estimation.

3.2×speedup
at 1,000 frames
0.2143depth RMSE
at 1,000 frames
0.976normal consistency
on NRGBD
0extra training
or token pruning

Depth Estimation on ScanNet-50

FramesMethodAbsRel ↓RMSE ↓RMSE-log ↓δ1 ↑δ3 ↑
100VGGT0.07070.25760.13530.92890.9917
FastVGGT0.06910.25040.13310.93080.9923
VGGT-R (Ours)0.06640.23800.12640.93700.9929
500VGGT0.06980.25250.13340.92940.9923
FastVGGT0.06650.25350.12930.93460.9932
VGGT-R (Ours)0.06620.23330.12430.93840.9934
Best results are highlighted. VGGT-R remains stable from 20 to 1,000 input frames and further improves as more views become available.

Visualization

Qualitative comparison of dense depth, reconstruction quality and register-guided token scoring.

Depth Estimation

VGGT-R depth estimation results compared with VGGT and ground truth

VGGT-R reduces depth errors around object boundaries and challenging regions while producing more coherent geometry.

3D Point Cloud Reconstruction

Point cloud example onePoint cloud example two

Dense reconstructions retain fine-grained spatial information without explicit patch-token pruning.

Register-Guided Token Selection

RGB input
RGB
Register relevance
Relevance
Local variation
Local variation
Temporal novelty
Temporal novelty
Final score
Final score

The three cues capture complementary semantic, structural and temporal information.

BibTeX

@misc{jiang2027vggtr,
  title  = {VGGT-R: You Can Really Accelerate Vision Geometry Grounded
            Transformer with the Underestimated Registers},
  author = {Jiang, Xuefeng and Geng, Zibin and Ma, Yuan and Li, Jia
            and Jia, Peijin and Sun, Sheng and Huang, Wenke and others},
  year   = {2027}
}