CPPF++: Uncertainty-Aware Sim2Real Object Pose Estimation by Vote Aggregation

June 12, 2026 · View on GitHub

CPPF++: Uncertainty-Aware Sim2Real Object Pose Estimation by Vote Aggregation

IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) 2024

TPAMI Project Page

Yang You, Wenhao He, Jin Liu, Hongkai Xiong, Weiming Wang, Cewu Lu

Overview

CPPF++ is a novel method for category-level 6D object pose estimation that bridges the sim-to-real gap using only 3D CAD models for training. Building upon the point-pair voting scheme of CPPF, our approach introduces a probabilistic framework that significantly improves robustness and accuracy in real-world scenarios.

Key Innovations

  • Uncertainty-aware voting: Models probabilistic distribution of point pairs to handle vote collisions
  • N-point tuple features: Enhanced contextual information through multi-point feature aggregation
  • Sim2Real robustness: Training exclusively on synthetic CAD models, generalizing to real images
  • DiversePose 300: New challenging dataset for category-level pose estimation

Performance Highlights

  • No real pose annotations required for training
  • State-of-the-art sim2real performance on NOCS REAL275
  • Superior generalization to novel datasets (Wild6D, PhoCAL, DiversePose)
  • Real-time inference capability for practical applications

📋 Table of Contents

🎬 Results in the Wild

Casual video captured on iPhone 15 - no smoothing, no object mesh required teaser

Check our new object pose benchmark PACE on ECCV 2024.

Also, check our work RPMArt on IROS 2024 that uses CPPF++ to get the pose of articulated objects (e.g., microwaves, drawers).

🚀 Quick Start

Prerequisites

  • Python 3.8+
  • PyTorch 1.8+
  • CUDA 11.0+
  • CMake 3.16+

Installation

  1. Clone the repository

    git clone https://github.com/qq456cvb/CPPF2.git
    cd CPPF2
    
  2. Install Python dependencies

    pip install torch torchvision torchaudio numpy opencv-python imageio tqdm omegaconf
    # Additional dependencies for the complete pipeline:
    pip install torch-scatter visdom matplotlib pillow fire
    # For optimization (optional):
    pip install lietorch
    
  3. Build SHOT descriptor extraction

    cd src_shot
    mkdir build && cd build
    cmake .. -A x64 -DCMAKE_BUILD_TYPE=Release
    cmake --build . --config Release
    

Quick Demo

Try our method on your own videos to generate pose estimation GIFs:

python demo.py --input_video your_video.mp4 --intrinsics_path camera_intrinsics.npy --object_category mug --output_gif result.gif

Required parameters:

  • --input_video: Path to input video file (e.g., .mp4, .avi)
  • --intrinsics_path: Path to camera intrinsics numpy file (.npy)
  • --object_category: Target object category (e.g., mug, bottle, bowl)

Optional parameters:

  • --output_gif: Output GIF filename (default: output.gif)
  • --input_depth_dir: Directory containing depth files (optional, for better accuracy)
  • --frame_skip: Process every N-th frame (default: 1, for faster processing use higher values)
  • --fps: GIF frame rate (default: 10)

Camera Intrinsics Format: The camera intrinsics should be a 3x3 numpy array saved as .npy file:

import numpy as np
# Example camera intrinsics matrix
intrinsics = np.array([
    [fx,  0, cx],
    [ 0, fy, cy], 
    [ 0,  0,  1]
])
np.save('camera_intrinsics.npy', intrinsics)

Where fx, fy are focal lengths and cx, cy are principal point coordinates.

📊 Evaluation

Pre-trained Models

Pre-trained checkpoints for the NOCS categories are included in the ckpts/ directory of this repository:

  • DINO models (ckpts/dino/): Visual feature extraction
  • SHOT models (ckpts/shot/): Geometric feature extraction

Evaluate on NOCS REAL275

  1. Download the NOCS REAL275 test images from the official source and extract them so the frames live under NOCS/real_test/.
  2. Download the Mask-RCNN detection results provided by SAR-Net from Google Drive and extract the real_test folder (containing results_*.pkl) to data/mrcnn_mask_results/real_test.
  3. Run the evaluation:
python eval.py
# or point to a custom mask location:
python eval.py --mask_dir /path/to/mrcnn_mask_results/real_test

Predictions are written to NOCS/nocs_output/ and mAP curves are printed at the end.

🔧 Training

Our method uses an ensemble of DINO (visual) and SHOT (geometric) features for optimal performance.

1. Prepare Training Data

You can create a conda environment with all training dependencies from the provided spec:

conda env create -f environment.yml
conda activate cppf2

Training data is rendered on the fly from ShapeNet v2 CAD models, so you need to download ShapeNetCore.v2 first (only the six NOCS categories are used: bottle, bowl, camera, can, laptop, mug). Place (or symlink) it at data/ShapeNetCore.v2, or point shapenet_root in config/config.yaml to its location:

# config/config.yaml
shapenet_root: /path/to/ShapeNetCore.v2

Then render the synthetic views and dump training features (RGB-D renderings, DINO descriptors, SHOT features) for all six categories:

python dataset.py

This writes ~100 samples per object to data/category_training_data/<category_id>/ and may take several hours. It requires a GPU (for DINOv2 feature extraction) and an EGL-capable machine (for headless pyrender rendering).

2. Train Individual Models

# Train DINO model (visual features)
python train_dino.py

# Train SHOT model (geometric features)  
python train_shot.py

3. Custom Object Training

We provide a comprehensive tutorial for training on your own objects:

  • 📓 Interactive Tutorial: train_custom.ipynb

For detailed instructions, see the Custom Training Section below.

📚 Training with Custom Objects

📓 Interactive Tutorial: Check train_custom.ipynb for a complete walkthrough with examples!

🗂️ Datasets

Supported Datasets

  • NOCS REAL275: Category-level pose estimation benchmark
  • YCB-Video: Instance-level object pose dataset
  • Wild6D: In-the-wild pose estimation
  • PhoCAL: Photorealistic object dataset
  • DiversePose 300: Our new diverse pose dataset, download from Hugging Face or Google Drive

Data Processing Scripts

We provide utilities to convert datasets into REAL275 format:

python data/wild6d_convert2real275.py
python data/phocal_convert2real275.py

🏗️ Method Overview

Architecture

CPPF++ extends the point-pair voting framework with several key innovations:

  1. Probabilistic Voting: Instead of deterministic votes, we model the uncertainty of each point pair's contribution
  2. N-point Tuples: Enhanced feature representation using multiple point relationships
  3. Ensemble Learning: Combining DINO (visual) and SHOT (geometric) features
  4. Uncertainty Estimation: Explicit modeling of pose estimation confidence
  • CPPF: Our previous point-pair voting framework
  • RPMArt: Extension to articulated objects (IROS 2024)
  • SAR-Net: Source for Mask-RCNN masks used in evaluation

📈 Updates & Changelog

  • 2024/06/27 - 📓 Added custom object training tutorial (train_custom.ipynb)
  • 2024/06/16 - 🎬 Demo code release for reproducing results
  • 2024/06/15 - 🎉 Paper accepted to IEEE TPAMI!
  • 2024/04/10 - 🔧 Data processing scripts for Wild6D and PhoCAL datasets
  • 2024/04/01 - 💾 Pre-trained models available in ckpts/
  • 2024/03/28 - 🚀 Major performance improvements across all datasets
  • 2023/09/06 - 📊 Significant method updates with enhanced NOCS and YCB-Video performance

🤝 Contributing

We welcome contributions! Please feel free to:

  • 🐛 Report issues or bugs
  • 💡 Suggest new features or improvements
  • 📖 Improve documentation
  • 🔧 Submit pull requests

📧 Contact

For questions about the paper or code, please:

  • Open a GitHub issue for technical problems
  • Contact Yang You for research discussions

📄 Citation

If you find our work helpful, please consider citing:

@ARTICLE{youcppf2,
  author={You, Yang and He, Wenhao and Liu, Jin and Xiong, Hongkai and Wang, Weiming and Lu, Cewu},
  journal={IEEE Transactions on Pattern Analysis and Machine Intelligence}, 
  title={CPPF++: Uncertainty-Aware Sim2Real Object Pose Estimation by Vote Aggregation}, 
  year={2024},
  volume={46},
  number={12},
  pages={9239-9254},
  keywords={Pose estimation;Training;Solid modeling;Noise measurement;Uncertainty;Training data;Filtering;Object pose estimation;sim-to-real transfer;point-pair features;dataset creation},
  doi={10.1109/TPAMI.2024.3419038}
}

📜 License

This project is released under the MIT License. See LICENSE for details.