ICLR 2025

3DCorrEnhance

Multiview Equivariance Improves 3D Correspondence Understanding with Minimal Feature Finetuning

1Stanford University2University of Southern California
Overview

Do vision transformers understand 3D?

Foundation models such as DINOv2 are trained on 2D images, yet 3D tasks increasingly rely on their features. We measure how view-equivariant those features are, meaning whether the same 3D point gets the same feature from every viewpoint. Equivariance turns out to predict performance on pose estimation, tracking and semantic correspondence, and a light finetuning on 3D correspondences improves it for every model we tried.

1.8B

correspondence pairs to measure equivariance, from 1,000 Objaverse objects in 42 views

+9.58

3cm-3deg pose accuracy on OnePose-LowTex for DINOv2 after finetuning

+6.45

Average Jaccard on TAP-Vid-DAVIS tracking, and +5.06 PCK@0.05 on PF-PASCAL

1object

one multiview pair and a single iteration already give about half of the full gain or more

3D equivariance

The same 3D point should get the same feature

We render 1,000 Objaverse objects from 42 views (42,000 images and 1.8 billion dense correspondences) and sample 1,000 real MVImgNet objects (33.3 million correspondences). For each point in one view, we find its nearest neighbour in feature space in another view, and score how far it lands from the true match: APE, the average pixel error, and PCDP, the percentage of correct dense points. We compare DINOv2, DINOv2 with registers, MAE, CLIP and DeiT.

An Objaverse horse skeleton from a reference view and two other views, with the features of DINOv2, DINOv2-Reg, MAE, CLIP and DeiT shown as PCA colours; DINOv2's colours stay the most consistent across views.
Features as PCA colours across three views. MAE mixes up the head and the body, CLIP and DeiT change the chest's features between views, and DINOv2 is the most consistent.

Equivariance predicts 3D task performance

We test three tasks that rely on correspondences, from easy to hard. Across models, lower APE (more equivariant features) goes with better scores on all of them.

One-shot pose estimation: 2D descriptors from reference views are back-projected into a database and matched to a query image, then RANSAC-PnP gives the pose.

One-shot pose estimation

Same object, rigid motion. Match a query image to stored views, then solve RANSAC-PnP. OnePose-LowTex and YCB-Video.

Video tracking: a point in the first frame is found in later frames by nearest-neighbour search on dense features.

Video tracking

Same object, non-rigid motion. Follow points through a video by feature similarity. TAP-Vid-DAVIS.

Semantic transfer: keypoints on a reference airplane are transferred to another airplane by nearest-neighbour search on features.

Semantic correspondence

Different objects of a category, any viewpoint. Transfer keypoints by feature similarity. PF-PASCAL.

Scatter plot: pose accuracy on OnePose-LowTex against APE for five ViTs; DINOv2 has the lowest APE and the highest accuracy.
Scatter plot: tracking Average Jaccard on TAP-Vid-DAVIS against APE for five ViTs; lower APE goes with higher AJ.
Scatter plot: average recall on YCB-Video against APE for five ViTs.
Scatter plot: PF-PASCAL keypoint matching accuracy against APE for five ViTs.

The points line up from the top left to the bottom right: more equivariant features, better 3D task performance. DINOv2 leads on both axes.

Finetuning

Pull corresponding pixels together

If equivariance helps, can we train for it? We finetune each model so that pixels showing the same 3D point in two views get similar features.

  1. Two viewsPick an Objaverse object and two of its rendered views, and sample pixels that see the same 3D points.
  2. Light adaptersLoRA on the last four transformer blocks, plus one 3×3 convolution that mixes neighbouring patches before upsampling.
  3. Ranking lossSmoothAP makes each pixel's true match rank first among candidates. It beats a contrastive loss with a fixed margin.
  4. Short training10K iterations with AdamW at learning rate 1e-5, on a 10K-object subset of Objaverse.
Overview: feature PCA colours of an acorn and a stapler from a pretrained ViT and from the same ViT finetuned on one synthetic object; radar charts show finetuned DINOv2, MAE, CLIP and DeiT beating the originals on pose estimation, tracking and semantic transfer.
Finetuned on one synthetic object, a ViT gives more consistent features on other objects (left) and better scores on pose estimation, tracking and semantic transfer (right).

Before and after, on unseen objects and scenes

Three input views
DINOv2 features before finetuning DINOv2 features after finetuning
BeforeAfter finetuning

DINOv2 features as PCA colours, before and after finetuning. Drag across the image: after finetuning, features are smoother, less noisy and more consistent between views. Objects from MVImgNet; scene from TAP-Vid-DAVIS.

Results

Every model gets better at every 3D task

The same finetuning recipe improves DINOv2, DINOv2 with registers, MAE, CLIP and DeiT, and it is trained only on synthetic objects, yet it also improves equivariance on real MVImgNet images. Pick a benchmark and a metric.

PretrainedFinetuned

All numbers from the paper and its appendix. The main metric of every benchmark improves for every model; a few secondary ones barely move (for example tracking occlusion accuracy for MAE and CLIP). "More backbones" covers DINOv2 Small to Giant, with and without registers, and ConvNeXt, on the three main benchmarks.

One object

A little 3D goes a long way

Finetuning on a single object already brings large improvements, and which object hardly matters: six random Objaverse objects all work, even an untextured hemisphere. Training on one multiview pair for a single iteration improves DINOv2 by 4.85 (3cm-3deg pose), 3.55 (tracking AJ) and 3.47 (PCK@0.05).

Bar charts of DINOv2 finetuned on six different single objects, including an untextured hemisphere; all beat the pretrained model, shown as dashed lines.
Finetuning on six different single objects. Dashed lines: the pretrained model.
An untextured hemisphere from four viewpoints, with DINOv2 features before and after finetuning; after finetuning, edges and inward and outward views get consistent features.
An untextured hemisphere: after finetuning, edges and inward and outward views get consistent features.
One-shot pose accuracy on OnePose-LowTex against the number of training iterations on one object; the curve jumps after the first iteration.
Tracking metrics on TAP-Vid-DAVIS against the number of training iterations on one object.
PF-PASCAL PCK against the number of training iterations on one object.
Performance against training iterations on a single object (0 to 10,000). The jump happens at the first iteration.
In the wild

Drop-in better features

Methods that use DINO features can simply swap in the finetuned ones. In Wild Gaussians, which uses DINOv2 to handle occluders in novel view synthesis, the finetuned features improve PSNR, SSIM and LPIPS on almost every scene. In LERF, they sharpen the language-embedded relevancy maps.

Novel view synthesis in the wild with Wild Gaussians, swapping in the finetuned DINOv2 features. Bold: better.
FeaturesMountainFountainCornerPatioSpotPatio-High
PSNR ↑SSIM ↑LPIPS ↓PSNR ↑SSIM ↑LPIPS ↓PSNR ↑SSIM ↑LPIPS ↓PSNR ↑SSIM ↑LPIPS ↓PSNR ↑SSIM ↑LPIPS ↓PSNR ↑SSIM ↑LPIPS ↓
Wild Gaussians, original DINOv220.820.6680.23920.900.6680.21323.510.8100.15221.310.8020.13423.960.7770.16522.040.7340.202
Wild Gaussians, finetuned DINOv221.010.6720.23420.970.6720.21223.740.8100.15121.230.8020.13324.010.7780.16322.110.7340.201
LERF relevancy maps for a query: with the original DINO regularizer and with the finetuned one, which localizes the object more tightly.
LERF with the original DINO regularizer (middle) and with the finetuned one (right).
Mutual nearest-neighbour matches between an object with and without its background, using pretrained DINOv2 features: few matches.
Pretrained DINOv2: matches between an object with and without its background.
The same pair with finetuned features: many more matches.
Finetuned: more matches. Inliers on 1K COCO objects, before → after: DINOv2 99 → 159 · DINOv2-Reg 76 → 148 · MAE 97 → 196 · CLIP 18 → 61 · DeiT 25 → 81.

Finetuning targets 3D correspondence. On tasks that need other things, features stay close to the original: Paris-H instance retrieval 75.92 → 76.23, VOC2012 segmentation mIoU 83.60 → 82.65, NYUv2 depth δ1 86.88 → 85.48 (DINOv2-Base, from the appendix).

Design choices

What matters in the recipe

Number of convolution layers added after the ViT. DINOv2-Base. Bold: best in each column; highlighted: the setting used.
VariantOnePose-LowTexTAP-Vid-DAVISPF-PASCAL, diff. views
1cm-1deg3cm-3deg5cm-5degAJδavgOAPCK@0.05PCK@0.10PCK@0.15
0 conv layers11.6953.8572.8344.5060.7984.0844.8257.1465.26
1 conv layer13.5858.0377.3546.8563.8484.1547.2560.7667.57
2 conv layers13.1256.1475.4547.4263.2584.1246.3258.0564.90
3 conv layers12.1553.6374.4646.8462.1482.9041.6053.9760.22
Training loss. DINOv2-Base. Bold: best in each column; highlighted: the setting used.
VariantOnePose-LowTexTAP-Vid-DAVISPF-PASCAL, diff. views
1cm-1deg3cm-3deg5cm-5degAJδavgOAPCK@0.05PCK@0.10PCK@0.15
SmoothAP13.5858.0377.3546.8563.8484.1547.2560.7667.57
Contrastive13.2855.5775.6843.7962.2081.8446.7058.0866.21
Differentiable Procrustes12.9255.0074.8643.6061.3282.7443.8957.2264.66
Finetuning data. DINOv2-Base. Bold: best in each column; highlighted: the setting used.
VariantOnePose-LowTexTAP-Vid-DAVISPF-PASCAL, diff. views
1cm-1deg3cm-3deg5cm-5degAJδavgOAPCK@0.05PCK@0.10PCK@0.15
Objaverse (synthetic objects)13.5858.0377.3546.8563.8484.1547.2560.7667.57
MVImgNet (real objects)13.6556.9874.6141.5358.8982.6745.1357.9365.40
Scene-centric (RealEstate10K, Spaces, LLFF)15.9560.7976.3547.3663.0780.2741.7352.3360.33
Learning rate. DINOv2-Base. Bold: best in each column; highlighted: the setting used.
VariantOnePose-LowTexTAP-Vid-DAVISPF-PASCAL, diff. views
1cm-1deg3cm-3deg5cm-5degAJδavgOAPCK@0.05PCK@0.10PCK@0.15
learning rate 1e-611.8655.0373.1244.7962.5683.1747.3460.1068.23
learning rate 3e-613.0557.4575.8945.9363.3283.7347.2060.5067.21
learning rate 1e-513.5858.0377.3546.8563.8484.1547.2560.7667.57
learning rate 3e-513.1558.3377.4946.7063.4583.3545.7057.9665.99

One convolution layer works best; SmoothAP beats contrastive and Procrustes losses; object-centric Objaverse beats real MVImgNet and scene-centric data overall; results are stable across learning rates.

Code & demo

Use the features

  • Demo: try the finetuned features on your own images in the Hugging Face Space.
  • Models: finetuned DINOv2 Small, Base, Large and Giant, and the other ViTs, on Hugging Face.
  • Code: finetuning on Objaverse and evaluation on pose estimation, tracking and semantic transfer.
from finetune import FinetuneDINO   # from the GitHub repo

model = FinetuneDINO.load_from_checkpoint(
    'https://huggingface.co/qq456cvb/3DCorrEnhance/resolve/main/dinov2_base.ckpt',
    r=4, backbone_size='base').eval().cuda()

# rgb: 3 x H x W float tensor in [0, 1]
with torch.no_grad():
    feats = model.get_feature_wo_kp(rgb[None].cuda(), normalize=True)[0]   # H x W x F
python evaluate.py --ckpt /path/to/ckpt --pose --tracking --transfer
Abstract
Read the full abstract

Vision foundation models, particularly the ViT family, have revolutionized image understanding by providing rich semantic features. However, despite their success in 2D comprehension, their abilities on grasping 3D spatial relationships are still unclear. In this work, we evaluate and enhance the 3D awareness of ViT-based models. We begin by systematically assessing their ability to learn 3D equivariant features, specifically examining the consistency of semantic embeddings across different viewpoints. Our findings indicate that improved 3D equivariance leads to better performance on various downstream tasks, including pose estimation, tracking, and semantic transfer. Building on this insight, we propose a simple yet effective finetuning strategy based on 3D correspondences, which significantly enhances the 3D correspondence understanding of existing vision models. Remarkably, finetuning on a single object for one iteration results in substantial gains.

Cite

BibTeX

@inproceedings{you2025multiview,
  title={Multiview Equivariance Improves 3D Correspondence Understanding with Minimal Feature Finetuning},
  author={You, Yang and Li, Yixin and Deng, Congyue and Wang, Yue and Guibas, Leonidas},
  booktitle={International Conference on Learning Representations (ICLR)},
  year={2025},
  url={https://openreview.net/forum?id=CNO4rbSV6v}
}