CVPR 2020

KeypointNet

A Large-scale 3D Keypoint Dataset Aggregated from Numerous Human Annotations

Yang You* Yujing Lou* Chengkun Li* Zhoujun Cheng Liangwei Li Lizhuang Ma Cewu Lu Weiming Wang†
Shanghai Jiao Tong University*Equal contribution†Corresponding author
Overview

Keypoints that many people agree on

KeypointNet is a large-scale, diverse 3D keypoint dataset built on ShapeNetCore. Instead of one expert's template, many annotators clicked the points they found important, and a new aggregation method turned their noisy, independent clicks into keypoints with semantic ids that are consistent across every model of a category. We also benchmark deep networks and hand-crafted detectors on two tasks: keypoint saliency and keypoint correspondence.

8,234

3D models from ShapeNetCore

103K+

semantic keypoints (103,447 in the release)

16

object categories, from airplanes to mugs

4–24

keypoints per model, with symmetry groups

Explore

Same colour, same keypoint

Every keypoint has a semantic id shared by all models of its category, so the same colour marks the same part on every object. Drag to rotate all nine models together, and hover a keypoint to find its counterparts.

Category

airplane

Semantic ids

Shuffle streams random models of the category from the Hugging Face dataset: 2,048-point clouds as released.

Why KeypointNet

Large, template-free and in correspondence

No template

Annotators choose the keypoints

Earlier 3D datasets label a fixed template designed by an expert, which is biased and often incomplete. Here, many annotators each chose their own keypoints.

Consistent ids

Keypoints correspond

Each keypoint has a semantic id shared across the category, so the data supports both keypoint detection and semantic correspondence.

16 categories

General objects, at scale

The only earlier template-free dataset in the table below has 43 models. KeypointNet has 8,234, from 16 categories.

How KeypointNet compares

Comparison of 3D keypoint datasets, from the paper. Correspondence: keypoints are indexed consistently across models. Template-free: no hard-coded keypoint template.
DatasetDomainCorrespondenceTemplate-freeInstancesCategoriesKeypointsFormat
FAUSThuman✓✗1001689Kmesh
SyncSpecCNNchair✓✗6,2431~60Kpoint cloud
Dutagaci et al.general✗✓4316<1Kmesh
Kim et al.general✓✗4044~3Kmesh
PASCAL 3D+general✗✗36,29212150K+RGB with 3D model
KeypointNet (ours)general✓✓8,23416103K+point cloud & mesh
How it was built

From noisy clicks to consistent keypoints

Different people click different points, and their clicks are never exact. KeypointNet aggregates them by minimizing a fidelity loss: it learns an embedding of every point together with a set of ground-truth keypoints that stay close, in that embedding, to what people annotated.

Keypoint aggregation pipeline on caps: raw annotations, embedding space, fidelity error map, potential keypoints after NMS, t-SNE clustering, and aggregated keypoints after human verification.
The aggregation pipeline, from raw annotations (top left) to aggregated keypoints (bottom left).
  1. AnnotateIn a web tool, each annotator clicks up to 24 points per model. Keypoints must be shared by the category, spread over the object and distinct in meaning.
  2. EmbedA PointConv network maps every point to an embedding. Each click is modelled as a Gaussian, because people make small mistakes.
  3. Fidelity errorFor every point, the embedding distances to the matching clicks on the other models of the category are summed into an error map.
  4. Non-minimum suppressionLocal minima of the error map become candidate keypoints on each model.
  5. ClusterCandidates are projected to 2D with t-SNE and grouped by embedding, which gives the semantic ids shared across models.
  6. VerifyExperts check the result with simple priors such as rotational symmetry and centrosymmetry, and fix keypoints the method missed.

Why not cluster the clicks directly? Clustering needs a distance threshold, cannot separate closely spaced keypoints such as the four on an airplane's tail, and gives no labels that match across models. Distances in the learned embedding are also less sensitive to mis-clicks than geodesic distances.

The annotation tool

Annotators work in a browser. The model to label is on the right; models they have already labelled stay on the left, which helps them keep their own keypoint indices consistent.

Keypoints are clicked on meshes and then transferred to point clouds of 2,048 points, so both forms are released.

Screenshot of the web annotation tool: five labelled airplanes with a keypoint on the nose on the left, the next airplane on the right, and keypoint indices 01 to 20 in a column.
The web annotation tool.
Statistics

What's inside

Counts per category, from the released annotations. Splits are train, validation and test at 7:1:2. Click a category to open it in the explorer.

Models, keypoints, semantic ids and keypoints per model are counted from the released annotations/all.json; splits and the number of annotators per category are from the paper.
CategoryModelsKeypointsSemantic idsPer modelTrain / val / testAnnotators
1,022
13,830
195–17715 / 102 / 20526
492
7,880
248–24344 / 49 / 9915
146
1,841
208–16102 / 14 / 308
380
6,260
184–18266 / 38 / 769
38
226
65–626 / 4 / 86
1,002
21,403
2214–22701 / 100 / 20117
999
12,060
218–17699 / 100 / 20015
697
5,900
94–9487 / 70 / 14013
90
776
95–962 / 10 / 188
270
1,358
64–6189 / 27 / 545
439
2,634
66–6307 / 44 / 8812
298
3,878
147–14208 / 30 / 607
186
2,043
1110–11130 / 18 / 389
141
1,323
106–1098 / 14 / 299
1,124
9,047
127–12786 / 113 / 22517
910
12,988
245–21637 / 91 / 18218
Total
8,234
103,447
–4–245,757 / 824 / 1,653–
Benchmarks

Semantic keypoints are hard to detect

We benchmark eight deep networks and three hand-crafted detectors on two tasks. Keypoint saliency asks which points are keypoints. Keypoint correspondence also asks which semantic id each one has. All numbers use a strict distance threshold of 0.01.

≤ 0.9

Hand-crafted detectors miss them

Harris3D, SIFT3D and ISS3D reach at most 0.9 average mIoU: their interest points are geometric, not semantic.

28.5 · 40.9

Best saliency, still low

DGCNN has the best average saliency of the 8 networks: 28.5 mIoU and 40.9 mAP.

61.7%

Correspondence: best PCK

RSNet has the best average PCK and is best on 8 of 16 categories, but no method gives fully consistent keypoints at 0.01.

5 of 8

Small categories are hardest

On caps, the smallest category (38 models), 5 of the 8 networks score 0.0 saliency mIoU.

Predict whether each point is a keypoint, without semantic labels. Scored by mean IoU, and by mean average precision for methods that output keypoint probabilities.

Keypoint saliency mIoU (%) at distance threshold 0.01, per category.
CategoryDeep networksHand-crafted
PointNetPointNet++RSNetSpiderCNNPointConvRSCNNDGCNNGraphCNNHarris3DSIFT3DISS3D
airplane10.433.533.724.640.031.139.20.71.61.21.1
bathtub0.818.523.113.029.020.425.517.00.31.01.4
bed3.016.428.416.425.126.232.422.20.90.50.9
bottle0.522.129.43.633.625.640.323.30.51.11.8
cap0.013.226.70.00.07.00.00.00.00.40.6
car7.324.030.411.835.925.427.023.00.91.01.7
chair7.718.222.718.225.522.024.90.60.70.70.7
guitar0.124.330.57.428.027.331.917.01.20.70.7
helmet0.07.015.80.07.78.917.19.70.00.40.8
knife0.019.021.714.926.621.131.518.71.90.60.4
laptop23.630.939.825.643.135.244.833.30.30.20.1
motorcycle0.121.931.512.035.923.630.30.60.70.80.8
mug0.216.324.12.219.815.923.90.50.00.50.9
skateboard0.013.721.61.89.321.922.611.71.00.70.8
table20.929.339.727.944.127.441.024.50.40.30.3
vessel7.116.718.312.921.718.124.415.11.21.01.1
Average5.120.327.312.026.622.328.513.60.70.70.9

Bold: best in each row, as marked in the paper. Hand-crafted detectors output no probabilities, so they have no mAP.

Download & use

Get the data

  • annotations/: one JSON file per category, plus all.json. Each keypoint has its 3D position, semantic id, point index in the point cloud, mesh face and barycentric coordinates, and symmetry groups.
  • pcds/: coloured point clouds of 2,048 points, one .pcd per model.
  • ShapeNetCore.v2.ply/: coloured triangle meshes, with vertex colours from the diffuse textures.
  • Code: train/val/test splits, visualization and benchmark scripts on GitHub.
  • License: MIT. Models come from ShapeNetCore.
# annotations and point clouds (about 650 MB)
pip install -U "huggingface_hub[cli]"
hf download qq456cvb/KeypointNet --repo-type dataset \
  --local-dir KeypointNet \
  --include "annotations/*" --include "pcds/*"
import json, numpy as np

anns = json.load(open("KeypointNet/annotations/chair.json"))
a = anns[0]
pts = np.loadtxt(
    f"KeypointNet/pcds/{a['class_id']}/{a['model_id']}.pcd",
    skiprows=10)[:, :3]                   # (2048, 3) xyz
kps = [(k["pcd_info"]["point_index"], k["semantic_id"])
       for k in a["keypoints"]]
Abstract
Read the full abstract

Detecting 3D objects keypoints is of great interest to the areas of both graphics and computer vision. There have been several 2D and 3D keypoint datasets aiming to address this problem in a data-driven way. These datasets, however, either lack scalability or bring ambiguity to the definition of keypoints. Therefore, we present KeypointNet: the first large-scale and diverse 3D keypoint dataset that contains 103,450 keypoints and 8,234 3D models from 16 object categories, by leveraging numerous human annotations. To handle the inconsistency between annotations from different people, we propose a novel method to aggregate these keypoints automatically, through minimization of a fidelity loss. Finally, ten state-of-the-art methods are benchmarked on our proposed dataset.

Cite

BibTeX

@inproceedings{you2020keypointnet,
  title={KeypointNet: A Large-Scale 3D Keypoint Dataset Aggregated From Numerous Human Annotations},
  author={You, Yang and Lou, Yujing and Li, Chengkun and Cheng, Zhoujun and Li, Liangwei and Ma, Lizhuang and Lu, Cewu and Wang, Weiming},
  booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
  pages={13647--13656},
  year={2020}
}