Official implementation for "QS-Attn: Query-Selected Attention for Contrastive Learning in I2I Translation" (CVPR 2022)

Last update: Dec 16, 2022

Overview

QS-Attn: Query-Selected Attention for Contrastive Learning in I2I Translation (CVPR2022)

Unpaired image-to-image (I2I) translation often requires to maximize the mutual information between the source and the translated images across different domains, which is critical for the generator to keep the source content and prevent it from unnecessary modifications. The self-supervised contrastive learning has already been successfully applied in the I2I. By constraining features from the same location to be closer than those from different ones, it implicitly ensures the result to take content from the source. However, previous work uses the features from random locations to impose the constraint, which may not be appropriate since some locations contain less information of source domain. Moreover, the feature itself does not reflect the relation with others. This paper deals with these problems by intentionally selecting significant anchor points for contrastive learning. We design a query-selected attention (QS-Attn) module, which compares feature distances in the source domain, giving an attention matrix with a probability distribution in each row. Then we select queries according to their measurement of significance, computed from the distribution. The selected ones are regarded as anchors for contrastive loss. At the same time, the reduced attention matrix is employed to route features in both domains, so that source relations maintain in the synthesis. We validate our proposed method in three different I2I datasets, showing that it increases the image quality without adding learnable parameters.

QS-Attn applies attention to select anchors for contrastive learning in single-direction I2I task

Getting Started

Prerequisites

Ubuntu 16.04
NVIDIA GPU + CUDA CuDNN
Python 3 Please use pip install -r requirements.txt to install the dependencies.

Pretrained Models

We provide Global, Local and Global+Local models for three datasets.

Model	Cityscapes	Horse2zebra	AFHQ
Global	Cityscapes_Global	Horse2zebra_Global	AFHQ_Global
Local	Cityscapes_Local	Horse2zebra_Local	AFHQ_Local
Global+Local	Cityscapes_Global+Local	Horse2zebra_Global+Local	AFHQ_Global+Local

Training

Download horse2zebra dataset :

bash ./datasets/download_qsattn_dataset.sh horse2zebra

Train the global model:

python train.py \
--dataroot=datasets/horse2zebra \
--name=horse2zebra_global \
--QS_mode=global

You can use visdom to view the training loss: Run python -m visdom.server and click the URL http://localhost:8097.

Inference

Test the global model:

python test.py \
--dataroot=datasets/horse2zebra \
--name=horse2zebra_qsattn_global \
--QS_mode=global

Citation

If you use this code for your research, please cite

@article{hu2022qs,
  title={QS-Attn: Query-Selected Attention for Contrastive Learning in I2I Translation},
  author={Hu, Xueqi and Zhou, Xinyue and Huang, Qiusheng and Shi, Zhengyi and Sun, Li and Li, Qingli},
  journal={arXiv preprint arXiv:2203.08483},
  year={2022}
}

Official implementation for "QS-Attn: Query-Selected Attention for Contrastive Learning in I2I Translation" (CVPR 2022)

Related tags

Overview

QS-Attn: Query-Selected Attention for Contrastive Learning in I2I Translation (CVPR2022)

Getting Started

Prerequisites

Pretrained Models

Training

Inference

Citation

Owner

Xueqi Hu

[EMNLP 2021] MuVER: Improving First-Stage Entity Retrieval with Multi-View Entity Representations

Facial Image Inpainting with Semantic Control

Variational Attention: Propagating Domain-Specific Knowledge for Multi-Domain Learning in Crowd Counting (ICCV, 2021)

The official repo for CVPR2021——ViPNAS: Efficient Video Pose Estimation via Neural Architecture Search.

Code for the paper "Adversarial Generator-Encoder Networks"

The 1st Place Solution of the Facebook AI Image Similarity Challenge (ISC21) : Descriptor Track.

Monocular 3D pose estimation. OpenVINO. CPU inference or iGPU (OpenCL) inference.

A PyTorch implementation of " EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks."

PCAM: Product of Cross-Attention Matrices for Rigid Registration of Point Clouds

Code for Boundary-Aware Segmentation Network for Mobile and Web Applications

PyTorch implementation of CVPR'18 - Perturbative Neural Networks

A package for music online and offline rhythmic information analysis including music Beat, downbeat, tempo and meter tracking.

[ICCV 2021] FaPN: Feature-aligned Pyramid Network for Dense Image Prediction

Planner_backend - Academic planner application designed for students and counselors.

Optimizes image files by converting them to webp while also updating all references.

WORD: Revisiting Organs Segmentation in the Whole Abdominal Region

Anti-UAV base on PaddleDetection

pyhsmm - library for approximate unsupervised inference in Bayesian Hidden Markov Models (HMMs) and explicit-duration Hidden semi-Markov Models (HSMMs), focusing on the Bayesian Nonparametric extensions, the HDP-HMM and HDP-HSMM, mostly with weak-limit approximations.

Code for KDD'20 "An Efficient Neighborhood-based Interaction Model for Recommendation on Heterogeneous Graph"

Unofficial Tensorflow 2 implementation of the paper Implicit Neural Representations with Periodic Activation Functions