Object pose, in the clutter of the real world.
PACE (Pose Annotations in Cluttered Environments) is a large-scale benchmark for object pose estimation and tracking in cluttered real-world scenes: everyday objects, rigid and articulated, piled together, occluded and moved by hand. It covers instance-level and category-level pose estimation, model-based and model-free tracking, and comes with PACE-Sim, a photo-realistic synthetic set for training.
real RGB-D frames
pose annotations
videos, 183 frames each on average
in 43 categories, rigid and articulated
environments with different levels of occlusion
annotation error, measured against NOCS REAL275 ground truth
PACE-Sim: physically based renders for training instance-level and category-level models.
Methods that solve REAL275 fail on PACE
On categories that exist in both datasets, state-of-the-art category-level methods score above 97% AP at 15°/5 cm on NOCS REAL275 for bottles, bowls and cans, and 69 to 86% for mugs. On PACE, none of them gets above 8%.
Look inside the annotations
Annotations include depth, instance masks, NOCS maps and the 6D pose of every object. Pick one and drag across the image to compare it with the RGB frame.
RGB3D pose Annotation
Depth is shown with a colour map for visibility; black marks pixels without depth.
Ten environments
Objects were captured in ten environments, each with its own level and kind of occlusion.










Far from solved
We evaluate state-of-the-art methods on two tracks, pose estimation and pose tracking. Pose estimation assumes ground-truth object detection, so the numbers isolate pose quality.
A classic beats deep learning
On instance-level pose, PPF, a point pair feature method that needs no training, reaches 52.2 AR. The best learned method, SurfEmb, reaches 9.0.
Category-level pose is wide open
The best AP at 0–20° and 0–5 cm is 9.9% on rigid objects and 1.1% on articulated ones.
Scale and sim2real both hurt
Methods fit one instance or category well with real training data, but trained on all of PACE they drop by 57.2 to 78.5 points. Depth-based HS-Pose, SGPA and DualPoseNet fall to 1.4 or below when trained on synthetic data.
Tracking struggles too
The best model-based tracker (ICG) reaches 38.1% ADD(-S) on rigid objects and 10.1% on articulated ones. The best model-free tracker stays below 13% at 5°/5 cm.
| Method | Input | Detection | ARVSD | ARMSSD | ARMSPD | AR |
|---|---|---|---|---|---|---|
| PPF | D | G.T. | 53.4 | 48.1 | 55.2 | 52.2 |
| CosyPose | RGB | G.T. | 1.4 | 0.3 | 11.5 | 4.4 |
| SurfEmb | RGB | G.T. | 6.2 | 3.0 | 17.8 | 9.0 |
| GDRNPP | RGB-D | G.T. | 3.6 | 2.1 | 15.4 | 7.0 |
| Method | Detection | IoU25 | IoU50 | AP (%), averaged over a threshold range | |||||
|---|---|---|---|---|---|---|---|---|---|
| 0–20° | 0–60° | 0–5 cm | 0–15 cm | 0–20°, 0–5 cm | 0–60°, 0–15 cm | ||||
| NOCS | Mask-RCNN | 0.2 | 0.0 | 1.2 | 4.3 | 17.0 | 33.5 | 1.1 | 4.2 |
| HS-Pose | G.T. | 36.6 | 2.7 | 5.3 | 8.6 | 48.7 | 81.0 | 3.9 | 8.1 |
| SGPA | G.T. | 2.6 | 1.2 | 5.6 | 11.4 | 16.1 | 50.7 | 3.3 | 10.1 |
| DualPoseNet | G.T. | 24.2 | 0.1 | 5.5 | 8.1 | 18.6 | 63.4 | 1.8 | 6.5 |
| SAR-Net | G.T. | 22.8 | 0.2 | 5.3 | 8.8 | 48.6 | 79.8 | 3.7 | 8.2 |
| CPPF++ | G.T. | 44.5 | 4.4 | 15.2 | 27.3 | 35.3 | 74.0 | 9.9 | 24.9 |
| ANCSH | G.T. | – | – | – | – | – | – | – | – |
| Method | Detection | IoU25 | IoU50 | AP (%), averaged over a threshold range | |||||
|---|---|---|---|---|---|---|---|---|---|
| 0–20° | 0–60° | 0–5 cm | 0–15 cm | 0–20°, 0–5 cm | 0–60°, 0–15 cm | ||||
| NOCS | Mask-RCNN | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 |
| HS-Pose | G.T. | 0.0 | 0.0 | 0.0 | 0.3 | 34.1 | 75.6 | 0.0 | 0.3 |
| SGPA | G.T. | 0.0 | 0.0 | 1.1 | 7.3 | 6.8 | 17.5 | 0.9 | 7.2 |
| DualPoseNet | G.T. | 0.3 | 0.0 | 0.0 | 0.0 | 27.3 | 65.2 | 0.0 | 0.0 |
| SAR-Net | G.T. | 0.1 | 0.0 | 0.0 | 0.8 | 38.5 | 69.7 | 0.0 | 0.8 |
| CPPF++ | G.T. | 3.6 | 0.0 | 1.7 | 6.2 | 31.1 | 66.7 | 1.1 | 5.9 |
| ANCSH | G.T. | 0.0 | 0.0 | 0.0 | 0.2 | 18.6 | 50.4 | 0.0 | 0.2 |
| Method | Input | ADD | ADD-S | ADD(-S) |
|---|---|---|---|---|
| RBOT | RGB | 7.1 | 10.3 | 7.4 |
| ICG | RGB-D | 35.6 | 48.1 | 38.1 |
| Method | Input | ADD | ADD-S | ADD(-S) |
|---|---|---|---|---|
| RBOT | RGB | 0.5 | 0.8 | 0.5 |
| ICG | RGB-D | 10.1 | 15.2 | 10.1 |
| Method | Training-free | Input | 5°5cm (%) ↑ | IoU25 (%) ↑ | Rerr (°) ↓ | Terr (cm) ↓ |
|---|---|---|---|---|---|---|
| BundleTrack | ✓ | RGB | 6.4 | 9.1 | 3.2 | 2.6 |
| CAPTRA | ✗ | D | 12.9 | 45.8 | 19.2 | 2.2 |
| 6-PACK | ✗ | RGB-D | 9.2 | 23.1 | 17.7 | 2.1 |
| Method | Training-free | Input | 5°5cm (%) ↑ | IoU25 (%) ↑ | Rerr (°) ↓ | Terr (cm) ↓ |
|---|---|---|---|---|---|---|
| BundleTrack | ✓ | RGB | 11.2 | 14.1 | 5.5 | 0.8 |
| CAPTRA | ✗ | D | 4.4 | 20.6 | 40.9 | 1.5 |
| 6-PACK | ✗ | RGB-D | 3.9 | 16.7 | 33.6 | 1.2 |
| Training setting | Instance-level (AR %) | Category-level (mean AP %, 0–60°, 0–15 cm) | |||||||
|---|---|---|---|---|---|---|---|---|---|
| CosyPose | SurfEmb | GDRNPP | NOCS | HS-Pose | SGPA | DualPoseNet | SAR-Net | CPPF++ | |
| Real → real, one instance / category | 76.4 | 86.7 | 88.2 | 76.9 | 81.2 | 86.7 | 58.6 | 59.3 | 83.4 |
| Synthetic → real, one instance / category | 67.1−9.3 | 73.7−13.0 | 78.1−10.1 | 55.8−21.1 | 1.1−80.1 | 0.3−86.4 | 1.4−57.2 | 33.2−26.1 | 55.8−27.6 |
| Real → real, all instances / categories | 8.2−68.2 | 11.2−75.5 | 9.7−78.5 | 0.3−76.6 | 5.6−75.6 | 9.8−76.9 | 1.2−57.4 | 2.1−57.2 | 13.7−69.7 |
Bold: best in each column, as marked in the paper. Detection “G.T.” means ground-truth masks are given.
How PACE compares
PACE is the only dataset in this comparison with all six properties: CAD models, moving objects, occlusion, marker-free images, articulated objects and piled clutter.
| Dataset | Input | Categories | Objects | Videos | Images | Annotations | CAD models | Moving objects | Occlusion | Marker-free | Articulated parts | Piled objects |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| YCB-Video | RGBD | – | 21 | 12 | 20K | 99K | ✓ | ✗ | ✓ | ✓ | ✗ | ✓ |
| LINEMOD-O | RGBD | – | 8 | 1 | 1.2K | 9.2K | ✓ | ✗ | ✓ | ✗ | ✗ | ✓ |
| NAVI | RGBD | – | 36 | 324 | 10K | 10k | ✓ | ✗ | ✗ | ✓ | ✗ | ✗ |
| StereoObj-1M | RGBD | – | 18 | 182 | 393K | 1.5M | ✓ | ✗ | ✓ | ✗ | ✗ | ✓ |
| NOCS-REAL275 | RGBD | 6 | 42 | 18 | 8K | – | ✓ | ✗ | ✓ | ✗ | ✗ | ✗ |
| Wild6D | RGBD | 5 | 1722 | 5166 | 1.1M | 1.1M | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ |
| Objectron | RGB | 9 | 17k | 14k | 4M | 4M | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ |
| Scan2CAD | RGBD | 9 | 3K | 1506 | – | 14K | ✓ | ✗ | ✓ | ✓ | ✗ | ✗ |
| Pix3D | RGBD | 9 | 395 | – | 10K | 10K | ✓ | ✗ | ✗ | ✓ | ✗ | ✗ |
| HANDAL | RGB | 17 | 212 | 2K | 308K | 308K | ✓ | ✓ | ✓ | ✓ | ✗ | ✗ |
| HouseCat6D | RGBD | 10 | 192 | 41 | 24K | 160K | ✓ | ✗ | ✓ | ✓ | ✗ | ✓ |
| ROPE | RGBD | 149 | 581 | 363 | 332K | 1.5M | ✓ | ✗ | ✓ | ✓ | ✗ | ✗ |
| PACE | RGBD | 43 | 238 | 300 | 55K | 258K | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
What's inside





A scalable annotation pipeline

- ScanEvery object is scanned with an EinScan Pro 2X, aligned to a shared frame per category, and labelled with its symmetries. Articulated objects come from AKB48.
- CaptureThree calibrated Intel RealSense D415 cameras record each scene at 1280×720, from 0.5 to 1.5 m away.
- Track static objectsA marker drives automatic pose tracking, corrected by hand every 40 frames, and is then inpainted out of the images.
- Track moving objectsBundleTrack follows objects moved by hand, with poses corrected every 10 frames.
- MasksOcclusion-aware masks are rendered from the poses; hands are segmented with SAM and removed.
To check accuracy, we re-annotated Scene 1 of NOCS REAL275 with this pipeline: the average error against its ground truth was 0.9° in rotation and 2.3 mm in translation. The annotation tool is open source.
Get the data
- Format: BOP: RGB, depth, masks, NOCS maps and poses per scene, plus scanned meshes and evaluation point clouds.
- Splits: real data is split 20/80 into validation and test; PACE-Sim provides the training sets.
- Evaluation: instance-level (BOP toolkit) and category-level code, with baseline predictions.
- License: MIT, except some 3D models from Sketchfab (CC BY 4.0) and GrabCAD.
# full dataset is about 870 GB; this fetches the real
# test set and object models (about 73 GB)
pip install -U "huggingface_hub[cli]"
hf download qq456cvb/PACE --repo-type dataset \
--local-dir dataset/pace \
--include "test_chunk_*" --include "models*.tar.gz" \
--include "model_splits/*" --include "test_targets_bop19.json"
cd dataset/pace
cat test_chunk_* > test.tar.gz
for f in *.tar.gz; do tar -xzf "$f"; doneRead the full abstract
We introduce PACE (Pose Annotations in Cluttered Environments), a large-scale benchmark designed to advance the development and evaluation of pose estimation methods in cluttered scenarios. PACE provides a large-scale real-world benchmark for both instance-level and category-level settings. The benchmark consists of 55K frames with 258K annotations across 300 videos, covering 238 objects from 43 categories and featuring a mix of rigid and articulated items in cluttered scenes. To annotate the real-world data efficiently, we develop an innovative annotation system with a calibrated 3-camera setup. Additionally, we offer PACE-Sim, which contains 100K photo-realistic simulated frames with 2.4M annotations across 931 objects. We test state-of-the-art algorithms in PACE along two tracks: pose estimation, and object pose tracking, revealing the benchmark's challenges and research opportunities.
BibTeX
@inproceedings{you2024pace,
title={PACE: A Large-Scale Dataset with Pose Annotations in Cluttered Environments},
author={You, Yang and Xiong, Kai and Yang, Zhening and Huang, Zhengxiang and Zhou, Junwei and Shi, Ruoxi and Fang, Zhou and Harley, Adam W. and Guibas, Leonidas and Lu, Cewu},
booktitle={European Conference on Computer Vision (ECCV)},
year={2024},
organization={Springer}
}