This code is an unofficial implementation of HiFiSinger.

Last update: Dec 23, 2022

Related tags

Overview

HiFiSinger

This code is an unofficial implementation of HiFiSinger. The algorithm is based on the following papers:

Chen, J., Tan, X., Luan, J., Qin, T., & Liu, T. Y. (2020). HiFiSinger: Towards High-Fidelity Neural Singing Voice Synthesis. arXiv preprint arXiv:2009.01776.
Ren, Y., Ruan, Y., Tan, X., Qin, T., Zhao, S., Zhao, Z., & Liu, T. Y. (2019). Fastspeech: Fast, robust and controllable text to speech. Advances in Neural Information Processing Systems, 32, 3171-3180.
Yamamoto, R., Song, E., & Kim, J. M. (2020, May). Parallel WaveGAN: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 6199-6203). IEEE.

Requirements

Please see the 'requirements.txt'.

Structure

Generator

In training, length regulator use target duration.

Discriminator

HiFiSinger uses Sub Frequency GAN(SF-GAN).
The frequency range of sampling is fixed and length range is randomized.

Used dataset

Code verification was conducted through a limited-sized, private Korean dataset.
- Thus, current Pattern_Generator.py and Datasets.py are based on the Korean.
Please report the information about any available open source dataset.
- The set of midi files with syncronized lyric and high resolution vocal wave files

Hyper parameters

Before proceeding, please set the pattern, inference, and checkpoint paths in 'Hyper_Parameters.yaml' according to your environment.

Sound
- Setting basic sound parameters.
Tokens
- The number of Lyric token.
Max_Note
- The highest note value for embedding.
Min/Max duration
- Mel length which model use.
- Min duration is used at pattern generating only.
Encoder
- Setting the encoder.
Duration_Predictor
- Setting for duration predictor
Decoder
- Setting for decoder.
Discriminator
- Setting for discriminator
- In frequency range, frequency is the index of mel dimension.
  - The index must be equal or less than Sould.Mel_Dim.
Vocoder_Path
- Setting the traced vocoder path.
- To generate this, please check Here
Train
- Setting the parameters of training.
Use_Mixed_Precision
- Setting mix precision usage.
- Need a Nvidia-Apex.
Inference_Batch_Size
- Setting the batch size when inference
Inference_Path
- Setting the inference path
Checkpoint_Path
- Setting the checkpoint path
Log_Path
- Setting the tensorboard log path
Device
- Setting which GPU device is used in multi-GPU enviornment.
- Or, if using only CPU, please set '-1'. (But, I don't recommend while training.)

Generate pattern

There is no available open source dataset.

Inference file path while training for verification.

Inference_for_Training
- There are two examples for inference.
- It is midi file based script.

Run

Command

python Train.py -s

-hp
- The hyper paramter file path
- This is required.
-s
- The resume step parameter.
- Default is 0.

This code is an unofficial implementation of HiFiSinger.

Related tags

Overview

HiFiSinger

Requirements

Structure

Generator

Discriminator

Used dataset

Hyper parameters

Generate pattern

Inference file path while training for verification.

Run

Command

Owner

Heejo You

A curated list of the top 10 computer vision papers in 2021 with video demos, articles, code and paper reference.

ResNEsts and DenseNEsts: Block-based DNN Models with Improved Representation Guarantees

ManiSkill-Learn is a framework for training agents on SAPIEN Open-Source Manipulation Skill Challenge (ManiSkill Challenge), a large-scale learning-from-demonstrations benchmark for object manipulation.

Official PyTorch implementation for FastDPM, a fast sampling algorithm for diffusion probabilistic models

A PyTorch Implementation of Single Shot Scale-invariant Face Detector.

RGB-D Local Implicit Function for Depth Completion of Transparent Objects

Official Repo for ICCV2021 Paper: Learning to Regress Bodies from Images using Differentiable Semantic Rendering

End-to-end face detection, cropping, norm estimation, and landmark detection in a single onnx model

The Self-Supervised Learner can be used to train a classifier with fewer labeled examples needed using self-supervised learning.

The code is an implementation of Feedback Convolutional Neural Network for Visual Localization and Segmentation.

This repository contains the code for the paper in EMNLP 2021: "HRKD: Hierarchical Relational Knowledge Distillation for Cross-domain Language Model Compression".

Examples of how to create colorful, annotated equations in Latex using Tikz.

TransReID: Transformer-based Object Re-Identification

Code Repository for Liquid Time-Constant Networks (LTCs)

Fast, general, and tested differentiable structured prediction in PyTorch

DSAC* for Visual Camera Re-Localization (RGB or RGB-D)

Object Detection using YOLO from PyImageSearch

Awesome Monocular 3D detection

Fre-GAN: Adversarial Frequency-consistent Audio Synthesis

This repository contains the code for the CVPR 2021 paper "GIRAFFE: Representing Scenes as Compositional Generative Neural Feature Fields"