Official pytorch implementation of the AAAI 2021 paper Semantic Grouping Network for Video Captioning

Last update: Nov 25, 2022

Related tags

Deep Learning SGN

Overview

Semantic Grouping Network for Video Captioning

Hobin Ryu, Sunghun Kang, Haeyong Kang, and Chang D. Yoo. AAAI 2021. [arxiv]

Environment

Ubuntu 16.04
CUDA 9.2
cuDNN 7.4.2
Java 8
Python 2.7.12
- PyTorch 1.1.0
- Other python packages specified in requirements.txt

Usage

1. Setup

$ pip install -r requirements.txt

2. Prepare Data

Download the GloVe Embedding from here and locate it at data/Embeddings/GloVe/GloVe_300.json.
Extract features from datasets and locate them at data/ /features/ .hdf5.

e.g. ResNet101 features of the MSVD dataset will be located at data/MSVD/features/ResNet101.hdf5.

I refer to this repo for extracting the ResNet101 features, and this repo for extracting the 3D-ResNext101 features.
Split the features into train, val, and test sets by running following commands.
```
$ python -m split.MSVD
$ python -m split.MSR-VTT
```

You can skip step 2-3 and download below files

MSVD
- ResNet-101 [train] [val] [test]
- 3D-ResNext-101 [train] [val] [test]
MSR-VTT
- ResNet-101 [train] [val] [test]
- 3D-ResNext-101 [train] [val] [test]

3. Prepare The Code for Evaluation

Clone the evaluation code from the official coco-evaluation repo.

$ git clone https://github.com/tylin/coco-caption.git
$ mv coco-caption/pycocoevalcap .
$ rm -rf coco-caption

4. Extract Negative Videos

$ python extract_negative_videos.py

or you can skip this step as the output files are already uploaded at data/ /metadata/neg_vids_ .json

5. Train

$ python train.py

You can change some hyperparameters by modifying config.py.

Pretrained Models - SGN(R101+RN)

*Disclaimer: The models above do not have the same weight as the models used in the paper (I trained them again because I lost).

6. Evaluate

$ python evaluate.py --ckpt_fpath

License

The source-code in this repository is released under MIT License.

Official pytorch implementation of the AAAI 2021 paper Semantic Grouping Network for Video Captioning

Related tags

Overview

Semantic Grouping Network for Video Captioning

Environment

Usage

1. Setup

2. Prepare Data

3. Prepare The Code for Evaluation

4. Extract Negative Videos

5. Train

6. Evaluate

License

Owner

Hobin Ryu

Code Repository for The Kaggle Book, Published by Packt Publishing

Server files for UltimateLabeling

Reducing Information Bottleneck for Weakly Supervised Semantic Segmentation (NeurIPS 2021)

Artificial intelligence technology inferring issues and logically supporting facts from raw text

VOLO: Vision Outlooker for Visual Recognition

TransMIL: Transformer based Correlated Multiple Instance Learning for Whole Slide Image Classification

Scaling Vision with Sparse Mixture of Experts

Using some basic methods to show linkages and transformations of robotic arms

Detectorch - detectron for PyTorch

Code for ACL2021 long paper: Knowledgeable or Educated Guess? Revisiting Language Models as Knowledge Bases

NL-Augmenter 🦎 → 🐍 A Collaborative Repository of Natural Language Transformations

A PyTorch implementation of NeRF (Neural Radiance Fields) that reproduces the results.

[CVPR 2021 Oral] Variational Relational Point Completion Network

This project hosts the code for implementing the ISAL algorithm for object detection and image classification

Weakly Supervised Text-to-SQL Parsing through Question Decomposition

Small little script to scrape, parse and check for active tor nodes. Can be used as proxies.

LSTC: Boosting Atomic Action Detection with Long-Short-Term Context

Code for "Neural Parts: Learning Expressive 3D Shape Abstractions with Invertible Neural Networks", CVPR 2021

RP-GAN: Stable GAN Training with Random Projections

Causal Influence Detection for Improving Efficiency in Reinforcement Learning