Code for Ditto: Building Digital Twins of Articulated Objects from Interaction

Last update: Dec 22, 2022

Overview

Ditto: Building Digital Twins of Articulated Objects from Interaction

Zhenyu Jiang, Cheng-Chun Hsu, Yuke Zhu

CVPR 2022, Oral

Project | arxiv

News

2022-04-28: We released the data generation code of Ditto here.

Introduction

Ditto (Digital Twins of Articulated Objects) is a model that reconstructs part-level geometry and articulation model of an articulated object given observations before and after an interaction. Specifically, we use a PointNet++ encoder to encoder the input point cloud observations, and fuse the subsampled point features with a simple attention layer. Then we use two independent decoders to propagate the fused point features into two sets of dense point features, for geometry reconstruction and articulation estimation separately. We construct feature grid/planes by projecting and pooling the point features, and query local features from the constructed feature grid/planes. Conditioning on local features, we use different decoders to predict occupancy, segmentation and joint parameters with respect to the query points. At then end, we can extract explicit geometry and articulation model from the implicit decoders.

If you find our work useful in your research, please consider citing.

Installation

Create a conda environment and install required packages.

conda env create -f conda_env_gpu.yaml -n Ditto

You can change the pytorch and cuda version in conda_env_gpu.yaml.

Build ConvONets dependents by running python scripts/convonet_setup.py build_ext --inplace.
Download the data, then unzip the data.zip under the repo's root.

Training

# single GPU
python run.py experiment=Ditto_s2m

# multiple GPUs
python run.py trainer.gpus=4 +trainer.accelerator='ddp' experiment=Ditto_s2m

# multiple GPUs + wandb logging
python run.py trainer.gpus=4 +trainer.accelerator='ddp' logger=wandb logger.wandb.group=s2m experiment=Ditto_s2m

Testing

# only support single GPU
python run_test.py experiment=Ditto_s2m trainer.resume_from_checkpoint=/path/to/trained/model/

Demo

Here is a minimum demo that starts from multiview depth maps before and after interaction and ends with a reconstructed digital twin. To run the demo, you need to install this library for visualization.

We provide the posed depth images of a real word laptop to run the demo. You can download from here and put it under data. You can also run demo your own data that follows the same format.

Data and pre-trained models

Data: here. Remeber to cite Shape2Motion and Abbatematteo et al. as well as Ditto when using these datasets.

Pre-trained models: Shape2Motion dataset, Synthetic dataset.

Useful tips

Run eval "$(python run.py -sc install=bash)" under the root directory, you can have auto-completion for commandline options.
Install pre-commit hooks by pip install pre-commit; pre-commit install, then you can have automatic formatting before each commit.

Related Repositories

Our code is based on this fantastic template Lightning-Hydra-Template.
We use ConvONets as our backbone.

Citing

@inproceedings{jiang2022ditto,
   title={Ditto: Building Digital Twins of Articulated Objects from Interaction},
   author={Jiang, Zhenyu and Hsu, Cheng-Chun and Zhu, Yuke},
   booktitle={arXiv preprint arXiv:2202.08227},
   year={2022}
}

Code for Ditto: Building Digital Twins of Articulated Objects from Interaction

Related tags

Overview

Ditto: Building Digital Twins of Articulated Objects from Interaction

News

Introduction

Installation

Training

Testing

Demo

Data and pre-trained models

Useful tips

Related Repositories

Citing

Owner

UT Robot Perception and Learning Lab

Price-Prediction-For-a-Dream-Home - A machine learning based linear regression trained model for house price prediction.

ProjectOxford-ClientSDK - This repo has moved :house: Visit our website for the latest SDKs & Samples

DeepSpeed is a deep learning optimization library that makes distributed training easy, efficient, and effective.

Latte: Cross-framework Python Package for Evaluation of Latent-based Generative Models

Vehicle detection using machine learning and computer vision techniques for Udacity's Self-Driving Car Engineer Nanodegree.

Unsupervised phone and word segmentation using dynamic programming on self-supervised VQ features.

Lightweight, Portable, Flexible Distributed/Mobile Deep Learning with Dynamic, Mutation-aware Dataflow Dep Scheduler; for Python, R, Julia, Scala, Go, Javascript and more

Dcf-game-infrastructure-public - Contains all the components necessary to run a DC finals (attack-defense CTF) game from OOO

MultiMix: Sparingly Supervised, Extreme Multitask Learning From Medical Images (ISBI 2021, MELBA 2021)

Code for models used in Bashiri et al., "A Flow-based latent state generative model of neural population responses to natural images".

Source code for the ACL-IJCNLP 2021 paper entitled "T-DNA: Taming Pre-trained Language Models with N-gram Representations for Low-Resource Domain Adaptation" by Shizhe Diao et al.

Sharpened cosine similarity torch - A Sharpened Cosine Similarity layer for PyTorch

Implementation of Geometric Vector Perceptron, a simple circuit for 3d rotation equivariance for learning over large biomolecules, in Pytorch. Idea proposed and accepted at ICLR 2021

A list of awesome PyTorch scholarship articles, guides, blogs, courses and other resources.

Conjugated Discrete Distributions for Distributional Reinforcement Learning (C2D)

A Simple LSTM-Based Solution for "Heartbeat Signal Classification and Prediction" in Tianchi

Model of an AI powered sign language interpreter.

UnivNet: A Neural Vocoder with Multi-Resolution Spectrogram Discriminators for High-Fidelity Waveform Generation

An implementation of MobileFormer

Code for: https://berkeleyautomation.github.io/bags/