So-ViT: Mind Visual Tokens for Vision Transformer

Last update: Nov 24, 2022

Related tags

Overview

So-ViT: Mind Visual Tokens for Vision Transformer

Introduction

This repository contains the source code under PyTorch framework and models trained on ImageNet-1K dataset for the following paper:

@articles{So-ViT,
    author = {Jiangtao Xie, Ruiren Zeng, Qilong Wang, Ziqi Zhou, Peihua Li},
    title = {So-ViT: Mind Visual Tokens for Vision Transformer},
    booktitle = {arXiv:2104.10935},
    year = {2021}
}

The Vision Transformer (ViT) heavily depends on pretraining using ultra large-scale datasets (e.g. ImageNet-21K or JFT-300M) to achieve high performance, while significantly underperforming on ImageNet-1K if trained from scratch. We propose a novel So-ViT model toward addressing this problem, by carefully considering the role of visual tokens.

Above all, for classification head, the ViT only exploits class token while entirely neglecting rich semantic information inherent in high-level visual tokens. Therefore, we propose a new classification paradigm, where the second-order, cross-covariance pooling of visual tokens is combined with class token for final classification. Meanwhile, a fast singular value power normalization is proposed for improving the second-order pooling.

Second, the ViT employs the naïve method of one linear projection of fixed-size image patches for visual token embedding, lacking the ability to model translation equivariance and locality. To alleviate this problem, we develop a light-weight, hierarchical module based on off-the-shelf convolutions for visual token embedding.

Classification results

Classification results (single crop 224x224, %) on ImageNet-1K validation set

Network	Top-1 Accuracy		Pre-trained models
Network	Paper reported	Upgrade	GoogleDrive	BaiduCloud
So-ViT-7	76.2	76.8	Coming soon	Coming soon
So-ViT-10	77.9	78.7	Coming soon	Coming soon
So-ViT-14	81.8	82.3	Coming soon	Coming soon
So-ViT-19	82.4	82.8	Coming soon	Coming soon

Installation and Usage

Install PyTorch (>=1.6.0)
Install timm (==0.3.4)
pip install thop
type git clone https://github.com/jiangtaoxie/So-ViT
prepare the dataset as follows

.
├── train
│   ├── class1
│   │   ├── class1_001.jpg
│   │   ├── class1_002.jpg
|   |   └── ...
│   ├── class2
│   ├── class3
│   ├── ...
│   ├── ...
│   └── classN
└── val
    ├── class1
    │   ├── class1_001.jpg
    │   ├── class1_002.jpg
    |   └── ...
    ├── class2
    ├── class3
    ├── ...
    ├── ...
    └── classN

for training from scracth

sh model_name.sh  # model_name = {So_vit_7/10/14/19}

Acknowledgment

pytorch: https://github.com/pytorch/pytorch

timm: https://github.com/rwightman/pytorch-image-models

T2T-ViT: https://github.com/yitu-opensource/T2T-ViT

Contact

If you have any questions or suggestions, please contact me

[email protected]

So-ViT: Mind Visual Tokens for Vision Transformer

Related tags

Overview

So-ViT: Mind Visual Tokens for Vision Transformer

Introduction

Classification results

Classification results (single crop 224x224, %) on ImageNet-1K validation set

Installation and Usage

for training from scracth

Acknowledgment

Contact

Owner

Jiangtao Xie

A set of tools for creating and testing machine learning features, with a scikit-learn compatible API

Repository for the paper : Meta-FDMixup: Cross-Domain Few-Shot Learning Guided byLabeled Target Data

NCVX (NonConVeX): A User-Friendly and Scalable Package for Nonconvex Optimization in Machine Learning.

TensorFlow implementation of Barlow Twins (Barlow Twins: Self-Supervised Learning via Redundancy Reduction)

High-quality single file implementation of Deep Reinforcement Learning algorithms with research-friendly features

When Does Pretraining Help? Assessing Self-Supervised Learning for Law and the CaseHOLD Dataset of 53,000+ Legal Holdings

Deep Distributed Control of Port-Hamiltonian Systems

This repository is maintained for the scientific paper tittled " Study of keyword extraction techniques for Electric Double Layer Capacitor domain using text similarity indexes: An experimental analysis "

A convolutional recurrent neural network for classifying A/B phases in EEG signals recorded for sleep analysis.

Process text, including tokenizing and representing sentences as vectors and Applying some concepts like RNN, LSTM and GRU to create a classifier can detect the language in which a sentence is written from among 17 languages.

[CVPR 2020] GAN Compression: Efficient Architectures for Interactive Conditional GANs

Spectral Tensor Train Parameterization of Deep Learning Layers

A minimal yet resourceful implementation of diffusion models (along with pretrained models + synthetic images for nine datasets)

Open source Python module for computer vision

Pytorch implementation of MaskGIT: Masked Generative Image Transformer

Official PyTorch Implementation of Learning Architectures for Binary Networks

Graph Analysis From Scratch

A 3D sparse LBM solver implemented using Taichi

Equivariant Imaging: Learning Beyond the Range Space

Graph Transformer Architecture. Source code for