Code release for "MERLOT Reserve: Neural Script Knowledge through Vision and Language and Sound"

Last update: Dec 11, 2022

Related tags

Overview

merlot_reserve

Code release for "MERLOT Reserve: Neural Script Knowledge through Vision and Language and Sound"

MERLOT Reserve (in submission) is a model for learning joint representations of vision, language, and sound from YouTube. The learned model can be used in a zero-shot or finetuned setting, where it does well on tasks like VCR and TVQA.

Visit our project page at rowanzellers.com/merlotreserve or read the full paper to learn more.

What's here

We are releasing the following:

JAX code, and model checkpoints, for the MERLOT model
Code for pretraining the model
Code for finetuning the model on VCR and TVQA
Code for doing zero-shot inference with the model

Environment and setup

There are two different ways to run MERLOT Reserve:

Pretraining on videos You'll need a TPU Pod VM for this. This step shouldn't be necessary for most people, as we have released model checkpoints.
Finetuning on VCR or TVQA I've done this on a TPU v3-8 VM. This should be possible on GPU(s), but I haven't tested this on such hardware.
Zero-shot inference I've ran this on a GPU (even an older, Titan X from 2016 works.)

Installation on a GPU Machine

Install Cuda 11.4 (I used this link) and CUDNN 8.2. You might have to add something like this to your PATH:

export LD_LIBRARY_PATH=/usr/local/cuda/lib64

Create the environment:

conda create --name mreserve python=3.8 && conda activate mreserve
conda install -y python=3.8 tqdm numpy pyyaml scipy ipython cython typing h5py pandas matplotlib

# Install jax
pip install jax[cuda11_cudnn82] -f https://storage.googleapis.com/jax-releases/jax_releases.html
# If doing this on TPUs instead of locally...
# pip install "jax[tpu]>=0.2.18" -f https://storage.googleapis.com/jax-releases/libtpu_releases.html

# This is needed sometimes https://stackoverflow.com/questions/66060487/valueerror-numpy-ndarray-size-changed-may-indicate-binary-incompatibility-exp
pip uninstall numpy
pip install numpy==1.19.5

pip install -r requirements.txt

You can then try out the interactive script at demo/demo_video.py. It will handle downloading the model checkpoint for you.

Installation on a Cloud TPU VM

See the instructions in pretrain/ to set up your environment on a TPU v3-8 VM.

Checkpoints

These should get auto-downloaded if you use PretrainedMerlotReserve in mreserve/modeling.py. All are flax checkpoint files:

# pretrained checkpoints
gs://merlotreserve/ckpts/base
gs://merlotreserve/ckpts/base_resadapt
gs://merlotreserve/ckpts/large
gs://merlotreserve/ckpts/large_resadapt

# finetuned checkpoints
gs://merlotreserve/vcr_ckpts/vcr_finetune_base
gs://merlotreserve/vcr_ckpts/vcr_finetune_large

gs://merlotreserve/tvqa_ckpts/tvqa_finetune_base
gs://merlotreserve/tvqa_ckpts/tvqa_finetune_large

# TVQA Data
gs://merlotreserve/finetune_data/tvqa/

# VCR data
gs://merlotreserve/finetune_data/vcr/

Code release for "MERLOT Reserve: Neural Script Knowledge through Vision and Language and Sound"

Related tags

Overview

merlot_reserve

What's here

Environment and setup

Installation on a GPU Machine

Installation on a Cloud TPU VM

Checkpoints

Owner

Rowan Zellers

Data Consistency for Magnetic Resonance Imaging

Implementation and replication of ProGen, Language Modeling for Protein Generation, in Jax

MlTr: Multi-label Classification with Transformer

DeepFaceLive - Live Deep Fake in python, Real-time face swap for PC streaming or video calls

Unit-Convertor - Unit Convertor Built With Python

Extract MNIST handwritten digits dataset binary file into bmp images

Code and data of the EMNLP 2021 paper "Mind the Style of Text! Adversarial and Backdoor Attacks Based on Text Style Transfer"

A PyTorch Toolbox for Face Recognition

implementation of paper - You Only Learn One Representation: Unified Network for Multiple Tasks

ML-Decoder: Scalable and Versatile Classification Head

Pixel-wise segmentation on VOC2012 dataset using pytorch.

Detecting Human-Object Interactions with Object-Guided Cross-Modal Calibrated Semantics

Anchor-free Oriented Proposal Generator for Object Detection

Model serving at scale

Tensorflow 2 Object Detection API kurulumu, GPU desteği, custom model hazırlama

We propose a new method for effective shadow removal by regarding it as an exposure fusion problem.

Fast Learning of MNL Model From General Partial Rankings with Application to Network Formation Modeling

This is a code repository for paper OODformer: Out-Of-Distribution Detection Transformer

Official repository for "Exploiting Session Information in BERT-based Session-aware Sequential Recommendation", SIGIR 2022 short.

Gesture-controlled Video Game. Just swing your finger and play the game without touching your PC