Do vision transformers understand 3D?
Foundation models such as DINOv2 are trained on 2D images, yet 3D tasks increasingly rely on their features. We measure how view-equivariant those features are, meaning whether the same 3D point gets the same feature from every viewpoint. Equivariance turns out to predict performance on pose estimation, tracking and semantic correspondence, and a light finetuning on 3D correspondences improves it for every model we tried.
correspondence pairs to measure equivariance, from 1,000 Objaverse objects in 42 views
3cm-3deg pose accuracy on OnePose-LowTex for DINOv2 after finetuning
Average Jaccard on TAP-Vid-DAVIS tracking, and +5.06 PCK@0.05 on PF-PASCAL
one multiview pair and a single iteration already give about half of the full gain or more
The same 3D point should get the same feature
We render 1,000 Objaverse objects from 42 views (42,000 images and 1.8 billion dense correspondences) and sample 1,000 real MVImgNet objects (33.3 million correspondences). For each point in one view, we find its nearest neighbour in feature space in another view, and score how far it lands from the true match: APE, the average pixel error, and PCDP, the percentage of correct dense points. We compare DINOv2, DINOv2 with registers, MAE, CLIP and DeiT.

Equivariance predicts 3D task performance
We test three tasks that rely on correspondences, from easy to hard. Across models, lower APE (more equivariant features) goes with better scores on all of them.

One-shot pose estimation
Same object, rigid motion. Match a query image to stored views, then solve RANSAC-PnP. OnePose-LowTex and YCB-Video.
Video tracking
Same object, non-rigid motion. Follow points through a video by feature similarity. TAP-Vid-DAVIS.

Semantic correspondence
Different objects of a category, any viewpoint. Transfer keypoints by feature similarity. PF-PASCAL.




The points line up from the top left to the bottom right: more equivariant features, better 3D task performance. DINOv2 leads on both axes.
Pull corresponding pixels together
If equivariance helps, can we train for it? We finetune each model so that pixels showing the same 3D point in two views get similar features.
- Two viewsPick an Objaverse object and two of its rendered views, and sample pixels that see the same 3D points.
- Light adaptersLoRA on the last four transformer blocks, plus one 3×3 convolution that mixes neighbouring patches before upsampling.
- Ranking lossSmoothAP makes each pixel's true match rank first among candidates. It beats a contrastive loss with a fixed margin.
- Short training10K iterations with AdamW at learning rate 1e-5, on a 10K-object subset of Objaverse.

Before and after, on unseen objects and scenes

BeforeAfter finetuning DINOv2 features as PCA colours, before and after finetuning. Drag across the image: after finetuning, features are smoother, less noisy and more consistent between views. Objects from MVImgNet; scene from TAP-Vid-DAVIS.
Every model gets better at every 3D task
The same finetuning recipe improves DINOv2, DINOv2 with registers, MAE, CLIP and DeiT, and it is trained only on synthetic objects, yet it also improves equivariance on real MVImgNet images. Pick a benchmark and a metric.
All numbers from the paper and its appendix. The main metric of every benchmark improves for every model; a few secondary ones barely move (for example tracking occlusion accuracy for MAE and CLIP). "More backbones" covers DINOv2 Small to Giant, with and without registers, and ConvNeXt, on the three main benchmarks.
A little 3D goes a long way
Finetuning on a single object already brings large improvements, and which object hardly matters: six random Objaverse objects all work, even an untextured hemisphere. Training on one multiview pair for a single iteration improves DINOv2 by 4.85 (3cm-3deg pose), 3.55 (tracking AJ) and 3.47 (PCK@0.05).





Drop-in better features
Methods that use DINO features can simply swap in the finetuned ones. In Wild Gaussians, which uses DINOv2 to handle occluders in novel view synthesis, the finetuned features improve PSNR, SSIM and LPIPS on almost every scene. In LERF, they sharpen the language-embedded relevancy maps.
| Features | Mountain | Fountain | Corner | Patio | Spot | Patio-High | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| PSNR ↑ | SSIM ↑ | LPIPS ↓ | PSNR ↑ | SSIM ↑ | LPIPS ↓ | PSNR ↑ | SSIM ↑ | LPIPS ↓ | PSNR ↑ | SSIM ↑ | LPIPS ↓ | PSNR ↑ | SSIM ↑ | LPIPS ↓ | PSNR ↑ | SSIM ↑ | LPIPS ↓ | |
| Wild Gaussians, original DINOv2 | 20.82 | 0.668 | 0.239 | 20.90 | 0.668 | 0.213 | 23.51 | 0.810 | 0.152 | 21.31 | 0.802 | 0.134 | 23.96 | 0.777 | 0.165 | 22.04 | 0.734 | 0.202 |
| Wild Gaussians, finetuned DINOv2 | 21.01 | 0.672 | 0.234 | 20.97 | 0.672 | 0.212 | 23.74 | 0.810 | 0.151 | 21.23 | 0.802 | 0.133 | 24.01 | 0.778 | 0.163 | 22.11 | 0.734 | 0.201 |



Finetuning targets 3D correspondence. On tasks that need other things, features stay close to the original: Paris-H instance retrieval 75.92 → 76.23, VOC2012 segmentation mIoU 83.60 → 82.65, NYUv2 depth δ1 86.88 → 85.48 (DINOv2-Base, from the appendix).
What matters in the recipe
| Variant | OnePose-LowTex | TAP-Vid-DAVIS | PF-PASCAL, diff. views | ||||||
|---|---|---|---|---|---|---|---|---|---|
| 1cm-1deg | 3cm-3deg | 5cm-5deg | AJ | δavg | OA | PCK@0.05 | PCK@0.10 | PCK@0.15 | |
| 0 conv layers | 11.69 | 53.85 | 72.83 | 44.50 | 60.79 | 84.08 | 44.82 | 57.14 | 65.26 |
| 1 conv layer | 13.58 | 58.03 | 77.35 | 46.85 | 63.84 | 84.15 | 47.25 | 60.76 | 67.57 |
| 2 conv layers | 13.12 | 56.14 | 75.45 | 47.42 | 63.25 | 84.12 | 46.32 | 58.05 | 64.90 |
| 3 conv layers | 12.15 | 53.63 | 74.46 | 46.84 | 62.14 | 82.90 | 41.60 | 53.97 | 60.22 |
| Variant | OnePose-LowTex | TAP-Vid-DAVIS | PF-PASCAL, diff. views | ||||||
|---|---|---|---|---|---|---|---|---|---|
| 1cm-1deg | 3cm-3deg | 5cm-5deg | AJ | δavg | OA | PCK@0.05 | PCK@0.10 | PCK@0.15 | |
| SmoothAP | 13.58 | 58.03 | 77.35 | 46.85 | 63.84 | 84.15 | 47.25 | 60.76 | 67.57 |
| Contrastive | 13.28 | 55.57 | 75.68 | 43.79 | 62.20 | 81.84 | 46.70 | 58.08 | 66.21 |
| Differentiable Procrustes | 12.92 | 55.00 | 74.86 | 43.60 | 61.32 | 82.74 | 43.89 | 57.22 | 64.66 |
| Variant | OnePose-LowTex | TAP-Vid-DAVIS | PF-PASCAL, diff. views | ||||||
|---|---|---|---|---|---|---|---|---|---|
| 1cm-1deg | 3cm-3deg | 5cm-5deg | AJ | δavg | OA | PCK@0.05 | PCK@0.10 | PCK@0.15 | |
| Objaverse (synthetic objects) | 13.58 | 58.03 | 77.35 | 46.85 | 63.84 | 84.15 | 47.25 | 60.76 | 67.57 |
| MVImgNet (real objects) | 13.65 | 56.98 | 74.61 | 41.53 | 58.89 | 82.67 | 45.13 | 57.93 | 65.40 |
| Scene-centric (RealEstate10K, Spaces, LLFF) | 15.95 | 60.79 | 76.35 | 47.36 | 63.07 | 80.27 | 41.73 | 52.33 | 60.33 |
| Variant | OnePose-LowTex | TAP-Vid-DAVIS | PF-PASCAL, diff. views | ||||||
|---|---|---|---|---|---|---|---|---|---|
| 1cm-1deg | 3cm-3deg | 5cm-5deg | AJ | δavg | OA | PCK@0.05 | PCK@0.10 | PCK@0.15 | |
| learning rate 1e-6 | 11.86 | 55.03 | 73.12 | 44.79 | 62.56 | 83.17 | 47.34 | 60.10 | 68.23 |
| learning rate 3e-6 | 13.05 | 57.45 | 75.89 | 45.93 | 63.32 | 83.73 | 47.20 | 60.50 | 67.21 |
| learning rate 1e-5 | 13.58 | 58.03 | 77.35 | 46.85 | 63.84 | 84.15 | 47.25 | 60.76 | 67.57 |
| learning rate 3e-5 | 13.15 | 58.33 | 77.49 | 46.70 | 63.45 | 83.35 | 45.70 | 57.96 | 65.99 |
One convolution layer works best; SmoothAP beats contrastive and Procrustes losses; object-centric Objaverse beats real MVImgNet and scene-centric data overall; results are stable across learning rates.
Use the features
- Demo: try the finetuned features on your own images in the Hugging Face Space.
- Models: finetuned DINOv2 Small, Base, Large and Giant, and the other ViTs, on Hugging Face.
- Code: finetuning on Objaverse and evaluation on pose estimation, tracking and semantic transfer.
from finetune import FinetuneDINO # from the GitHub repo model = FinetuneDINO.load_from_checkpoint( 'https://huggingface.co/qq456cvb/3DCorrEnhance/resolve/main/dinov2_base.ckpt', r=4, backbone_size='base').eval().cuda() # rgb: 3 x H x W float tensor in [0, 1] with torch.no_grad(): feats = model.get_feature_wo_kp(rgb[None].cuda(), normalize=True)[0] # H x W x F
python evaluate.py --ckpt /path/to/ckpt --pose --tracking --transfer
Read the full abstract
Vision foundation models, particularly the ViT family, have revolutionized image understanding by providing rich semantic features. However, despite their success in 2D comprehension, their abilities on grasping 3D spatial relationships are still unclear. In this work, we evaluate and enhance the 3D awareness of ViT-based models. We begin by systematically assessing their ability to learn 3D equivariant features, specifically examining the consistency of semantic embeddings across different viewpoints. Our findings indicate that improved 3D equivariance leads to better performance on various downstream tasks, including pose estimation, tracking, and semantic transfer. Building on this insight, we propose a simple yet effective finetuning strategy based on 3D correspondences, which significantly enhances the 3D correspondence understanding of existing vision models. Remarkably, finetuning on a single object for one iteration results in substantial gains.
BibTeX
@inproceedings{you2025multiview,
title={Multiview Equivariance Improves 3D Correspondence Understanding with Minimal Feature Finetuning},
author={You, Yang and Li, Yixin and Deng, Congyue and Wang, Yue and Guibas, Leonidas},
booktitle={International Conference on Learning Representations (ICLR)},
year={2025},
url={https://openreview.net/forum?id=CNO4rbSV6v}
}