PyTorchCV: A PyTorch-Based Framework for Deep Learning in Computer Vision.

Last update: Sep 14, 2022

Related tags

Overview

PyTorchCV: A PyTorch-Based Framework for Deep Learning in Computer Vision

@misc{CV2018,
  author =       {Donny You ([email protected])},
  howpublished = {\url{https://github.com/donnyyou/PyTorchCV}},
  year =         {2018}
}

This repository provides source code for some deep learning based cv problems. We'll do our best to keep this repository up to date. If you do find a problem about this repository, please raise it as an issue. We will fix it immediately.

Implemented Papers

Image Classification
- VGG: Very Deep Convolutional Networks for Large-Scale Image Recognition
- ResNet: Deep Residual Learning for Image Recognition
- DenseNet: Densely Connected Convolutional Networks
- ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices
- ShuffleNet V2: Practical Guidelines for Ecient CNN Architecture Design
Semantic Segmentation
- DeepLabV3: Rethinking Atrous Convolution for Semantic Image Segmentation
- PSPNet: Pyramid Scene Parsing Network
- DenseASPP: DenseASPP for Semantic Segmentation in Street Scenes
Object Detection
- SSD: Single Shot MultiBox Detector
- Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
- YOLOv3: An Incremental Improvement
- FPN: Feature Pyramid Networks for Object Detection
Pose Estimation
- CPM: Convolutional Pose Machines
- OpenPose: Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields
Instance Segmentation
- Mask R-CNN

Performances with PyTorchCV

Image Classification

ResNet: Deep Residual Learning for Image Recognition

Semantic Segmentation

PSPNet: Pyramid Scene Parsing Network

Model	Backbone	Training data	Testing data	mIOU	Pixel Acc	Setting
PSPNet Origin	3x3-ResNet101	ADE20K train	ADE20K val	41.96	80.64	-
PSPNet Ours	7x7-ResNet101	ADE20K train	ADE20K val	44.18	80.91	PSPNet

Object Detection

SSD: Single Shot MultiBox Detector

Model	Backbone	Training data	Testing data	mAP	FPS	Setting
SSD-300 Origin	VGG16	VOC07+12 trainval	VOC07 test	0.772	-	-
SSD-300 Ours	VGG16	VOC07+12 trainval	VOC07 test	0.786	-	SSD300
SSD-512 Origin	VGG16	VOC07+12 trainval	VOC07 test	0.798	-	-
SSD-512 Ours	VGG16	VOC07+12 trainval	VOC07 test	0.808	-	SSD512

Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks

Model	Backbone	Training data	Testing data	mAP	FPS	Setting
Faster R-CNN Origin	VGG16	VOC07 trainval	VOC07 test	0.699	-	-
Faster R-CNN Ours	VGG16	VOC07 trainval	VOC07 test	0.706	-	Faster R-CNN

YOLOv3: An Incremental Improvement

Pose Estimation

OpenPose: Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields

Instance Segmentation

Mask R-CNN

Commands with PyTorchCV

Take PSPNet as an example. ("tag" could be any string, include an empty one.)

Training

cd scripts/seg/cityscapes/
bash run_fs_pspnet_cityscapes_seg.sh train tag

Resume Training

cd scripts/seg/cityscapes/
bash run_fs_pspnet_cityscapes_seg.sh train tag

Validate

cd scripts/seg/cityscapes/
bash run_fs_pspnet_cityscapes_seg.sh val tag

Testing:

cd scripts/seg/cityscapes/
bash run_fs_pspnet_cityscapes_seg.sh test tag

Examples with PyTorchCV

Example output of VGG19-OpenPose

PyTorchCV: A PyTorch-Based Framework for Deep Learning in Computer Vision.

Related tags

Overview

PyTorchCV: A PyTorch-Based Framework for Deep Learning in Computer Vision

Implemented Papers

Performances with PyTorchCV

Image Classification

Semantic Segmentation

Object Detection

Pose Estimation

Instance Segmentation

Commands with PyTorchCV

Examples with PyTorchCV

Owner

Donny You

Supervised Contrastive Learning for Downstream Optimized Sequence Representations

Official Tensorflow implementation of "M-LSD: Towards Light-weight and Real-time Line Segment Detection"

HODEmu, is both an executable and a python library that is based on Ragagnin 2021 in prep.

Official implementation for "Symbolic Learning to Optimize: Towards Interpretability and Scalability"

Hierarchical Few-Shot Generative Models

An onlinel learning to rank python codebase.

Lipstick ain't enough: Beyond Color-Matching for In-the-Wild Makeup Transfer (CVPR 2021)

Implementation for Curriculum DeepSDF

CLDF dataset derived from Robbeets et al.'s "Triangulation Supports Agricultural Spread" from 2021

ConvMAE: Masked Convolution Meets Masked Autoencoders

Watch faces morph into each other with StyleGAN 2, StyleGAN, and DCGAN!

A fast python implementation of Ray Tracing in One Weekend using python and Taichi

ReAct: Out-of-distribution Detection With Rectified Activations

A very simple tool for situations where optimization with onnx-simplifier would exceed the Protocol Buffers upper file size limit of 2GB, or simply to separate onnx files to any size you want.

Distilled coarse part of LoFTR adapted for compatibility with TensorRT and embedded divices

A Context-aware Visual Attention-based training pipeline for Object Detection from a Webpage screenshot!

PyTorch implementation of Decoupling Value and Policy for Generalization in Reinforcement Learning

Memory-Augmented Model Predictive Control

OCRA (Object-Centric Recurrent Attention) source code

AQP is a modular pipeline built to enable the comparison and testing of different quality metric configurations.