[CVPR 2021] Forecasting the panoptic segmentation of future video frames

Last update: Nov 29, 2022

Overview

Panoptic Segmentation Forecasting

Colin Graber, Grace Tsai, Michael Firman, Gabriel Brostow, Alexander Schwing - CVPR 2021

We propose to study the novel task of ‘panoptic segmentation forecasting’: given a set of observed frames, the goal is to forecast the panoptic segmentation for a set of unobserved frames. We also propose a first approach to forecasting future panoptic segmentations. In contrast to typical semantic forecasting, we model the motion of individual object instances and the background separately. This makes instance information persistent during forecasting, and allows us to understand the motion of each moving object.

⚙️ Setup

Dependencies

Python 3.7
PyTorch 1.5.1
pyyaml
pandas
h5py
opencv
tensorboard
tqdm
pytorch_scatter 2.0.5
cityscapesscripts (for evaluation)
Google Cloud SDK (for downloading data/models)

Install the code using the following command: pip install -e ./

Data

To run this code, the gtFine_trainvaltest dataset will need to be downloaded from the Cityscapes website into the data/ directory.
The remainder of the required data can be downloaded using the script download_data.sh. By default, everything is downloaded into the data/ directory.
Training the background model requires generating a version of the semantic segmentation annotations where foreground regions have been removed. This can be done by running the script scripts/preprocessing/remove_fg_from_gt.sh.
Training the foreground model requires additionally downloading a pretrained MaskRCNN model. This can be found at this link. This should be saved as pretrained_models/fg/mask_rcnn_pretrain.pkl.
Training the background model requires additionally downloading a pretrained HarDNet model. This can be found at this link. This should be saved as pretrained_models/bg/hardnet70_cityscapes_model.pkl.

Running our code

The scripts directory contains scripts which can be used to train and evaluate the foreground, background, and egomotion models. Specifically:

scripts/odom/run_odom_train.sh trains the egomotion prediction model.
scripts/odom/export_odom.sh exports the odometry predictions, which can then be used during evaluation by other models
scripts/bg/run_bg_train.sh trains the background prediction model.
scripts/bg/run_export_bg_val.sh exports predictions make by the background using input reprojected point clouds which come from using predicted egomotion.
scripts/fg/run_fg_train.sh trains the foreground prediction model.
scripts/fg/run_fg_eval_panoptic.sh produces final panoptic semgnetation predictions based on the trained foreground model and exported background predictions. This also uses predicted egomotion as input.

We provide our pretrained foreground, background, and egomotion prediction models. The data downloading script additionally downloads these models into the directory pretrained_models/

✏️ 📄 Citation

If you found our work relevant to yours, please consider citing our paper:

@inproceedings{graber-2021-panopticforecasting,
 title   = {Panoptic Segmentation Forecasting},
 author  = {Colin Graber and
            Grace Tsai and
            Michael Firman and
            Gabriel Brostow and
            Alexander Schwing},
 booktitle = {Computer Vision and Pattern Recognition ({CVPR})},
 year = {2021}
}

[CVPR 2021] Forecasting the panoptic segmentation of future video frames

Related tags

Overview

Panoptic Segmentation Forecasting

⚙️ Setup

Dependencies

Data

Running our code

✏️ 📄 Citation

👩‍⚖️ License

Owner

Niantic Labs

Cluster-GCN: An Efficient Algorithm for Training Deep and Large Graph Convolutional Networks

Bayesian Generative Adversarial Networks in Tensorflow

Conceptual 12M is a dataset containing (image-URL, caption) pairs collected for vision-and-language pre-training.

High dimensional black-box optimizer using Latent Action Monte Carlo Tree Search algorithm

Lighthouse: Predicting Lighting Volumes for Spatially-Coherent Illumination

Code for our WACV 2022 paper "Hyper-Convolution Networks for Biomedical Image Segmentation"

💛 Code and Dataset for our EMNLP 2021 paper: "Perspective-taking and Pragmatics for Generating Empathetic Responses Focused on Emotion Causes"

Cross-modal Deep Face Normals with Deactivable Skip Connections

Yet Another Reinforcement Learning Tutorial

Py-FEAT: Python Facial Expression Analysis Toolbox

BOVText: A Large-Scale, Multidimensional Multilingual Dataset for Video Text Spotting

Implementation for paper LadderNet: Multi-path networks based on U-Net for medical image segmentation

A PyTorch-based library for semi-supervised learning

Accommodating supervised learning algorithms for the historical prices of the world's favorite cryptocurrency and boosting it through LightGBM.

3D mesh stylization driven by a text input in PyTorch

Official implementation of the Neurips 2021 paper Searching Parameterized AP Loss for Object Detection.

Source code for the GPT-2 story generation models in the EMNLP 2020 paper "STORIUM: A Dataset and Evaluation Platform for Human-in-the-Loop Story Generation"

Just-Now - This Is Just Now Login Friendlist Cloner Tools

Python/Rust implementations and notes from Proofs Arguments and Zero Knowledge

ILVR: Conditioning Method for Denoising Diffusion Probabilistic Models (ICCV 2021 Oral)