Keypoints that many people agree on
KeypointNet is a large-scale, diverse 3D keypoint dataset built on ShapeNetCore. Instead of one expert's template, many annotators clicked the points they found important, and a new aggregation method turned their noisy, independent clicks into keypoints with semantic ids that are consistent across every model of a category. We also benchmark deep networks and hand-crafted detectors on two tasks: keypoint saliency and keypoint correspondence.
3D models from ShapeNetCore
semantic keypoints (103,447 in the release)
object categories, from airplanes to mugs
keypoints per model, with symmetry groups
Same colour, same keypoint
Every keypoint has a semantic id shared by all models of its category, so the same colour marks the same part on every object. Drag to rotate all nine models together, and hover a keypoint to find its counterparts.
Category
Semantic ids
Shuffle streams random models of the category from the Hugging Face dataset: 2,048-point clouds as released.
Large, template-free and in correspondence
Annotators choose the keypoints
Earlier 3D datasets label a fixed template designed by an expert, which is biased and often incomplete. Here, many annotators each chose their own keypoints.
Keypoints correspond
Each keypoint has a semantic id shared across the category, so the data supports both keypoint detection and semantic correspondence.
General objects, at scale
The only earlier template-free dataset in the table below has 43 models. KeypointNet has 8,234, from 16 categories.
How KeypointNet compares
| Dataset | Domain | Correspondence | Template-free | Instances | Categories | Keypoints | Format |
|---|---|---|---|---|---|---|---|
| FAUST | human | ✓ | ✗ | 100 | 1 | 689K | mesh |
| SyncSpecCNN | chair | ✓ | ✗ | 6,243 | 1 | ~60K | point cloud |
| Dutagaci et al. | general | ✗ | ✓ | 43 | 16 | <1K | mesh |
| Kim et al. | general | ✓ | ✗ | 404 | 4 | ~3K | mesh |
| PASCAL 3D+ | general | ✗ | ✗ | 36,292 | 12 | 150K+ | RGB with 3D model |
| KeypointNet (ours) | general | ✓ | ✓ | 8,234 | 16 | 103K+ | point cloud & mesh |
From noisy clicks to consistent keypoints
Different people click different points, and their clicks are never exact. KeypointNet aggregates them by minimizing a fidelity loss: it learns an embedding of every point together with a set of ground-truth keypoints that stay close, in that embedding, to what people annotated.

- AnnotateIn a web tool, each annotator clicks up to 24 points per model. Keypoints must be shared by the category, spread over the object and distinct in meaning.
- EmbedA PointConv network maps every point to an embedding. Each click is modelled as a Gaussian, because people make small mistakes.
- Fidelity errorFor every point, the embedding distances to the matching clicks on the other models of the category are summed into an error map.
- Non-minimum suppressionLocal minima of the error map become candidate keypoints on each model.
- ClusterCandidates are projected to 2D with t-SNE and grouped by embedding, which gives the semantic ids shared across models.
- VerifyExperts check the result with simple priors such as rotational symmetry and centrosymmetry, and fix keypoints the method missed.
Why not cluster the clicks directly? Clustering needs a distance threshold, cannot separate closely spaced keypoints such as the four on an airplane's tail, and gives no labels that match across models. Distances in the learned embedding are also less sensitive to mis-clicks than geodesic distances.
The annotation tool
Annotators work in a browser. The model to label is on the right; models they have already labelled stay on the left, which helps them keep their own keypoint indices consistent.
Keypoints are clicked on meshes and then transferred to point clouds of 2,048 points, so both forms are released.

What's inside
Counts per category, from the released annotations. Splits are train, validation and test at 7:1:2. Click a category to open it in the explorer.
| Category | Models | Keypoints | Semantic ids | Per model | Train / val / test | Annotators |
|---|---|---|---|---|---|---|
1,022 | 13,830 | 19 | 5–17 | 715 / 102 / 205 | 26 | |
492 | 7,880 | 24 | 8–24 | 344 / 49 / 99 | 15 | |
146 | 1,841 | 20 | 8–16 | 102 / 14 / 30 | 8 | |
380 | 6,260 | 18 | 4–18 | 266 / 38 / 76 | 9 | |
38 | 226 | 6 | 5–6 | 26 / 4 / 8 | 6 | |
1,002 | 21,403 | 22 | 14–22 | 701 / 100 / 201 | 17 | |
999 | 12,060 | 21 | 8–17 | 699 / 100 / 200 | 15 | |
697 | 5,900 | 9 | 4–9 | 487 / 70 / 140 | 13 | |
90 | 776 | 9 | 5–9 | 62 / 10 / 18 | 8 | |
270 | 1,358 | 6 | 4–6 | 189 / 27 / 54 | 5 | |
439 | 2,634 | 6 | 6–6 | 307 / 44 / 88 | 12 | |
298 | 3,878 | 14 | 7–14 | 208 / 30 / 60 | 7 | |
186 | 2,043 | 11 | 10–11 | 130 / 18 / 38 | 9 | |
141 | 1,323 | 10 | 6–10 | 98 / 14 / 29 | 9 | |
1,124 | 9,047 | 12 | 7–12 | 786 / 113 / 225 | 17 | |
910 | 12,988 | 24 | 5–21 | 637 / 91 / 182 | 18 | |
| Total | 8,234 | 103,447 | – | 4–24 | 5,757 / 824 / 1,653 | – |
Semantic keypoints are hard to detect
We benchmark eight deep networks and three hand-crafted detectors on two tasks. Keypoint saliency asks which points are keypoints. Keypoint correspondence also asks which semantic id each one has. All numbers use a strict distance threshold of 0.01.
Hand-crafted detectors miss them
Harris3D, SIFT3D and ISS3D reach at most 0.9 average mIoU: their interest points are geometric, not semantic.
Best saliency, still low
DGCNN has the best average saliency of the 8 networks: 28.5 mIoU and 40.9 mAP.
Correspondence: best PCK
RSNet has the best average PCK and is best on 8 of 16 categories, but no method gives fully consistent keypoints at 0.01.
Small categories are hardest
On caps, the smallest category (38 models), 5 of the 8 networks score 0.0 saliency mIoU.
Predict whether each point is a keypoint, without semantic labels. Scored by mean IoU, and by mean average precision for methods that output keypoint probabilities.
| Category | Deep networks | Hand-crafted | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| PointNet | PointNet++ | RSNet | SpiderCNN | PointConv | RSCNN | DGCNN | GraphCNN | Harris3D | SIFT3D | ISS3D | |
| airplane | 10.4 | 33.5 | 33.7 | 24.6 | 40.0 | 31.1 | 39.2 | 0.7 | 1.6 | 1.2 | 1.1 |
| bathtub | 0.8 | 18.5 | 23.1 | 13.0 | 29.0 | 20.4 | 25.5 | 17.0 | 0.3 | 1.0 | 1.4 |
| bed | 3.0 | 16.4 | 28.4 | 16.4 | 25.1 | 26.2 | 32.4 | 22.2 | 0.9 | 0.5 | 0.9 |
| bottle | 0.5 | 22.1 | 29.4 | 3.6 | 33.6 | 25.6 | 40.3 | 23.3 | 0.5 | 1.1 | 1.8 |
| cap | 0.0 | 13.2 | 26.7 | 0.0 | 0.0 | 7.0 | 0.0 | 0.0 | 0.0 | 0.4 | 0.6 |
| car | 7.3 | 24.0 | 30.4 | 11.8 | 35.9 | 25.4 | 27.0 | 23.0 | 0.9 | 1.0 | 1.7 |
| chair | 7.7 | 18.2 | 22.7 | 18.2 | 25.5 | 22.0 | 24.9 | 0.6 | 0.7 | 0.7 | 0.7 |
| guitar | 0.1 | 24.3 | 30.5 | 7.4 | 28.0 | 27.3 | 31.9 | 17.0 | 1.2 | 0.7 | 0.7 |
| helmet | 0.0 | 7.0 | 15.8 | 0.0 | 7.7 | 8.9 | 17.1 | 9.7 | 0.0 | 0.4 | 0.8 |
| knife | 0.0 | 19.0 | 21.7 | 14.9 | 26.6 | 21.1 | 31.5 | 18.7 | 1.9 | 0.6 | 0.4 |
| laptop | 23.6 | 30.9 | 39.8 | 25.6 | 43.1 | 35.2 | 44.8 | 33.3 | 0.3 | 0.2 | 0.1 |
| motorcycle | 0.1 | 21.9 | 31.5 | 12.0 | 35.9 | 23.6 | 30.3 | 0.6 | 0.7 | 0.8 | 0.8 |
| mug | 0.2 | 16.3 | 24.1 | 2.2 | 19.8 | 15.9 | 23.9 | 0.5 | 0.0 | 0.5 | 0.9 |
| skateboard | 0.0 | 13.7 | 21.6 | 1.8 | 9.3 | 21.9 | 22.6 | 11.7 | 1.0 | 0.7 | 0.8 |
| table | 20.9 | 29.3 | 39.7 | 27.9 | 44.1 | 27.4 | 41.0 | 24.5 | 0.4 | 0.3 | 0.3 |
| vessel | 7.1 | 16.7 | 18.3 | 12.9 | 21.7 | 18.1 | 24.4 | 15.1 | 1.2 | 1.0 | 1.1 |
| Average | 5.1 | 20.3 | 27.3 | 12.0 | 26.6 | 22.3 | 28.5 | 13.6 | 0.7 | 0.7 | 0.9 |
| Category | PointNet | PointNet++ | RSNet | SpiderCNN | PointConv | RSCNN | DGCNN | GraphCNN |
|---|---|---|---|---|---|---|---|---|
| airplane | 9.3 | 44.6 | 38.8 | 24.7 | 51.2 | 45.0 | 52.0 | 0.7 |
| bathtub | 6.3 | 32.3 | 36.0 | 16.4 | 44.7 | 34.9 | 40.6 | 23.7 |
| bed | 9.4 | 23.4 | 42.5 | 17.1 | 43.7 | 40.1 | 42.8 | 31.5 |
| bottle | 8.2 | 38.2 | 44.3 | 8.2 | 46.5 | 40.1 | 61.6 | 42.6 |
| cap | 0.3 | 12.1 | 26.9 | 0.5 | 7.0 | 8.9 | 2.0 | 4.1 |
| car | 9.2 | 41.0 | 42.3 | 17.2 | 55.6 | 44.7 | 48.6 | 36.5 |
| chair | 6.5 | 29.4 | 26.2 | 20.4 | 39.8 | 33.0 | 34.6 | 0.6 |
| guitar | 1.4 | 27.7 | 34.2 | 9.9 | 39.3 | 34.1 | 44.1 | 21.6 |
| helmet | 1.3 | 7.4 | 21.3 | 0.5 | 8.3 | 10.8 | 19.9 | 12.5 |
| knife | 0.9 | 22.1 | 31.5 | 18.7 | 38.1 | 27.2 | 39.9 | 23.3 |
| laptop | 32.0 | 60.0 | 62.0 | 40.0 | 63.5 | 58.8 | 69.6 | 49.7 |
| motorcycle | 4.0 | 36.0 | 43.2 | 13.3 | 50.6 | 37.0 | 41.4 | 0.7 |
| mug | 3.9 | 21.4 | 35.2 | 4.0 | 32.0 | 18.6 | 33.2 | 0.6 |
| skateboard | 2.9 | 16.1 | 25.9 | 3.3 | 10.1 | 30.1 | 29.8 | 12.8 |
| table | 23.1 | 47.4 | 54.6 | 40.5 | 64.0 | 49.8 | 59.0 | 34.7 |
| vessel | 9.1 | 26.9 | 22.3 | 14.6 | 31.8 | 28.9 | 35.6 | 19.3 |
| Average | 8.0 | 30.4 | 36.7 | 15.6 | 39.1 | 33.9 | 40.9 | 19.7 |
Predict a fixed set of keypoints together with their semantic ids. Scored by the percentage of correct keypoints (PCK).
| Category | PointNet | PointNet++ | RSNet | SpiderCNN | PointConv | RSCNN | DGCNN | GraphCNN |
|---|---|---|---|---|---|---|---|---|
| airplane | 65.4 | 64.1 | 68.9 | 54.9 | 70.3 | 66.8 | 66.8 | 41.8 |
| bathtub | 57.2 | 45.4 | 56.0 | 36.8 | 54.4 | 47.9 | 51.5 | 19.5 |
| bed | 55.4 | 44.2 | 69.2 | 43.9 | 58.6 | 52.4 | 56.3 | 35.4 |
| bottle | 78.9 | 10.2 | 78.2 | 50.4 | 64.6 | 59.4 | 8.9 | 32.1 |
| cap | 16.7 | 18.8 | 47.9 | 0.0 | 25.0 | 18.8 | 39.6 | 8.3 |
| car | 66.6 | 55.1 | 65.7 | 47.1 | 65.7 | 56.5 | 58.6 | 36.9 |
| chair | 43.2 | 38.0 | 49.4 | 36.9 | 51.1 | 45.0 | 46.0 | 27.0 |
| guitar | 65.4 | 55.4 | 63.7 | 37.7 | 62.6 | 60.0 | 62.4 | 28.9 |
| helmet | 42.6 | 23.6 | 33.8 | 5.6 | 16.2 | 21.3 | 36.1 | 8.8 |
| knife | 32.6 | 43.8 | 56.9 | 32.2 | 47.4 | 51.4 | 52.1 | 20.0 |
| laptop | 76.1 | 64.6 | 80.3 | 63.4 | 74.2 | 66.9 | 73.9 | 65.9 |
| motorcycle | 62.1 | 49.9 | 67.4 | 36.6 | 61.7 | 52.8 | 57.7 | 34.4 |
| mug | 60.7 | 36.4 | 61.7 | 16.4 | 51.2 | 29.5 | 51.4 | 15.6 |
| skateboard | 50.4 | 42.3 | 69.2 | 15.8 | 51.8 | 51.7 | 52.3 | 26.2 |
| table | 63.8 | 57.1 | 72.3 | 61.0 | 69.5 | 64.8 | 70.5 | 17.3 |
| vessel | 42.5 | 37.9 | 46.7 | 36.8 | 46.7 | 41.2 | 46.8 | 22.1 |
| Average | 55.0 | 42.9 | 61.7 | 36.0 | 54.4 | 49.2 | 51.9 | 27.5 |
The same metrics as the distance threshold grows from 0 to 0.1. RSNet leads on correspondence at every threshold.





Bold: best in each row, as marked in the paper. Hand-crafted detectors output no probabilities, so they have no mAP.
Get the data
annotations/: one JSON file per category, plusall.json. Each keypoint has its 3D position, semantic id, point index in the point cloud, mesh face and barycentric coordinates, and symmetry groups.pcds/: coloured point clouds of 2,048 points, one.pcdper model.ShapeNetCore.v2.ply/: coloured triangle meshes, with vertex colours from the diffuse textures.- Code: train/val/test splits, visualization and benchmark scripts on GitHub.
- License: MIT. Models come from ShapeNetCore.
# annotations and point clouds (about 650 MB)
pip install -U "huggingface_hub[cli]"
hf download qq456cvb/KeypointNet --repo-type dataset \
--local-dir KeypointNet \
--include "annotations/*" --include "pcds/*"import json, numpy as np
anns = json.load(open("KeypointNet/annotations/chair.json"))
a = anns[0]
pts = np.loadtxt(
f"KeypointNet/pcds/{a['class_id']}/{a['model_id']}.pcd",
skiprows=10)[:, :3] # (2048, 3) xyz
kps = [(k["pcd_info"]["point_index"], k["semantic_id"])
for k in a["keypoints"]]Read the full abstract
Detecting 3D objects keypoints is of great interest to the areas of both graphics and computer vision. There have been several 2D and 3D keypoint datasets aiming to address this problem in a data-driven way. These datasets, however, either lack scalability or bring ambiguity to the definition of keypoints. Therefore, we present KeypointNet: the first large-scale and diverse 3D keypoint dataset that contains 103,450 keypoints and 8,234 3D models from 16 object categories, by leveraging numerous human annotations. To handle the inconsistency between annotations from different people, we propose a novel method to aggregate these keypoints automatically, through minimization of a fidelity loss. Finally, ten state-of-the-art methods are benchmarked on our proposed dataset.
BibTeX
@inproceedings{you2020keypointnet,
title={KeypointNet: A Large-Scale 3D Keypoint Dataset Aggregated From Numerous Human Annotations},
author={You, Yang and Lou, Yujing and Li, Chengkun and Cheng, Zhoujun and Li, Liangwei and Ma, Lizhuang and Lu, Cewu and Wang, Weiming},
booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
pages={13647--13656},
year={2020}
}