ECCV 2024

PACE

A Large-Scale Dataset with Pose Annotations in Cluttered Environments

1Stanford University2Shanghai Jiao Tong University3Horizon Robotics4UC San Diego
Overview

Object pose, in the clutter of the real world.

PACE (Pose Annotations in Cluttered Environments) is a large-scale benchmark for object pose estimation and tracking in cluttered real-world scenes: everyday objects, rigid and articulated, piled together, occluded and moved by hand. It covers instance-level and category-level pose estimation, model-based and model-free tracking, and comes with PACE-Sim, a photo-realistic synthetic set for training.

55K

real RGB-D frames

258K

pose annotations

300

videos, 183 frames each on average

238objects

in 43 categories, rigid and articulated

10

environments with different levels of occlusion

0.9°/ 2.3 mm

annotation error, measured against NOCS REAL275 ground truth

100Kframes
2.4Mannotations
931objects

PACE-Sim: physically based renders for training instance-level and category-level models.

Why PACE

Methods that solve REAL275 fail on PACE

On categories that exist in both datasets, state-of-the-art category-level methods score above 97% AP at 15°/5 cm on NOCS REAL275 for bottles, bowls and cans, and 69 to 86% for mugs. On PACE, none of them gets above 8%.

NOCS REAL275PACEAP (%) at 15° / 5 cm
HS-Pose
bottle
99.8
0.4
bowl
99.8
0.0
can
99.5
0.8
mug
86.2
7.5
DualPoseNet
bottle
97.5
0.3
bowl
99.7
2.0
can
97.7
0.5
mug
68.7
1.6
SAR-Net
bottle
98.2
2.1
bowl
98.2
0.0
can
97.4
0.3
mug
70.6
0.0
Explore

Look inside the annotations

Annotations include depth, instance masks, NOCS maps and the 6D pose of every object. Pick one and drag across the image to compare it with the RGB frame.

RGB frame Annotation overlay
RGB3D pose

Annotation

Depth is shown with a colour map for visibility; black marks pixels without depth.

Ten environments

Objects were captured in ten environments, each with its own level and kind of occlusion.

Objects in a plastic basket
basket
Objects on top of a box
box top
Objects in a cabinet
cabinet
Objects on a carpet
carpet
Objects on a chair
chair
Objects on a tiled floor
tiled floor
Objects on a desk
desk 1
Objects on another desk
desk 2
Objects on stairs
stairs
Objects on a wooden floor
wooden floor
Benchmarks

Far from solved

We evaluate state-of-the-art methods on two tracks, pose estimation and pose tracking. Pose estimation assumes ground-truth object detection, so the numbers isolate pose quality.

52.2 vs 9.0

A classic beats deep learning

On instance-level pose, PPF, a point pair feature method that needs no training, reaches 52.2 AR. The best learned method, SurfEmb, reaches 9.0.

9.9% · 1.1%

Category-level pose is wide open

The best AP at 0–20° and 0–5 cm is 9.9% on rigid objects and 1.1% on articulated ones.

−57.2 to −78.5

Scale and sim2real both hurt

Methods fit one instance or category well with real training data, but trained on all of PACE they drop by 57.2 to 78.5 points. Depth-based HS-Pose, SGPA and DualPoseNet fall to 1.4 or below when trained on synthetic data.

38.1% · 12.9%

Tracking struggles too

The best model-based tracker (ICG) reaches 38.1% ADD(-S) on rigid objects and 10.1% on articulated ones. The best model-free tracker stays below 13% at 5°/5 cm.

Average recall (%) under the BOP protocol, with ground-truth detections. PPF uses no learned model; the others are trained on synthetic PBR data.
MethodInputDetectionARVSDARMSSDARMSPDAR
PPFDG.T.53.448.155.252.2
CosyPoseRGBG.T.1.40.311.54.4
SurfEmbRGBG.T.6.23.017.89.0
GDRNPPRGB-DG.T.3.62.115.47.0

Bold: best in each column, as marked in the paper. Detection “G.T.” means ground-truth masks are given.

Dataset

How PACE compares

PACE is the only dataset in this comparison with all six properties: CAD models, moving objects, occlusion, marker-free images, articulated objects and piled clutter.

Instance-level datasets (top), category-level datasets (bottom). Numbers as reported in the paper; “–” where not available.
DatasetInputCategoriesObjectsVideosImagesAnnotationsCAD modelsMoving objectsOcclusionMarker-freeArticulated partsPiled objects
YCB-VideoRGBD–211220K99K✓✗✓✓✗✓
LINEMOD-ORGBD–811.2K9.2K✓✗✓✗✗✓
NAVIRGBD–3632410K10k✓✗✗✓✗✗
StereoObj-1MRGBD–18182393K1.5M✓✗✓✗✗✓
NOCS-REAL275RGBD642188K–✓✗✓✗✗✗
Wild6DRGBD5172251661.1M1.1M✗✗✗✓✗✗
ObjectronRGB917k14k4M4M✗✗✗✓✗✗
Scan2CADRGBD93K1506–14K✓✗✓✓✗✗
Pix3DRGBD9395–10K10K✓✗✗✓✗✗
HANDALRGB172122K308K308K✓✓✓✓✗✗
HouseCat6DRGBD101924124K160K✓✗✓✓✗✓
ROPERGBD149581363332K1.5M✓✗✓✓✗✗
PACERGBD4323830055K258K✓✓✓✓✓✓

What's inside

Bar chart of pose annotation counts for the 43 categories, from toys (about 18,000) down to ramen package; articulated categories box, scissor, cutter and clip are marked in red.
Pose annotations per category. Articulated categories (box, scissors, cutter, clip) are marked in red.
Stacked bar chart of object instances per category in PACE and PACE-Sim.
Object instances per category in PACE (real) and PACE-Sim.
Histogram of object sizes, peaking near 0.15 to 0.2 m, up to about 0.6 m.
Object sizes (bounding-box diagonal); most are near 0.2 m.
Donut charts of azimuth and elevation viewing-angle distributions.
Viewing angles: azimuth and elevation.
Donut chart of occlusion levels: minor, moderate and severe.
Occlusion levels: minor, moderate and severe.
How it was built

A scalable annotation pipeline

Overview of the PACE annotation pipeline: object scanning and alignment, calibrated three-camera capture with markers, marker-based and BundleTrack-assisted pose annotation, marker inpainting and mask generation.
  1. ScanEvery object is scanned with an EinScan Pro 2X, aligned to a shared frame per category, and labelled with its symmetries. Articulated objects come from AKB48.
  2. CaptureThree calibrated Intel RealSense D415 cameras record each scene at 1280×720, from 0.5 to 1.5 m away.
  3. Track static objectsA marker drives automatic pose tracking, corrected by hand every 40 frames, and is then inpainted out of the images.
  4. Track moving objectsBundleTrack follows objects moved by hand, with poses corrected every 10 frames.
  5. MasksOcclusion-aware masks are rendered from the poses; hands are segmented with SAM and removed.

To check accuracy, we re-annotated Scene 1 of NOCS REAL275 with this pipeline: the average error against its ground truth was 0.9° in rotation and 2.3 mm in translation. The annotation tool is open source.

Download & use

Get the data

  • Format: BOP: RGB, depth, masks, NOCS maps and poses per scene, plus scanned meshes and evaluation point clouds.
  • Splits: real data is split 20/80 into validation and test; PACE-Sim provides the training sets.
  • Evaluation: instance-level (BOP toolkit) and category-level code, with baseline predictions.
  • License: MIT, except some 3D models from Sketchfab (CC BY 4.0) and GrabCAD.
# full dataset is about 870 GB; this fetches the real
# test set and object models (about 73 GB)
pip install -U "huggingface_hub[cli]"
hf download qq456cvb/PACE --repo-type dataset \
  --local-dir dataset/pace \
  --include "test_chunk_*" --include "models*.tar.gz" \
  --include "model_splits/*" --include "test_targets_bop19.json"

cd dataset/pace
cat test_chunk_* > test.tar.gz
for f in *.tar.gz; do tar -xzf "$f"; done
Abstract
Read the full abstract

We introduce PACE (Pose Annotations in Cluttered Environments), a large-scale benchmark designed to advance the development and evaluation of pose estimation methods in cluttered scenarios. PACE provides a large-scale real-world benchmark for both instance-level and category-level settings. The benchmark consists of 55K frames with 258K annotations across 300 videos, covering 238 objects from 43 categories and featuring a mix of rigid and articulated items in cluttered scenes. To annotate the real-world data efficiently, we develop an innovative annotation system with a calibrated 3-camera setup. Additionally, we offer PACE-Sim, which contains 100K photo-realistic simulated frames with 2.4M annotations across 931 objects. We test state-of-the-art algorithms in PACE along two tracks: pose estimation, and object pose tracking, revealing the benchmark's challenges and research opportunities.

Cite

BibTeX

@inproceedings{you2024pace,
  title={PACE: A Large-Scale Dataset with Pose Annotations in Cluttered Environments},
  author={You, Yang and Xiong, Kai and Yang, Zhening and Huang, Zhengxiang and Zhou, Junwei and Shi, Ruoxi and Fang, Zhou and Harley, Adam W. and Guibas, Leonidas and Lu, Cewu},
  booktitle={European Conference on Computer Vision (ECCV)},
  year={2024},
  organization={Springer}
}