Charsiu: A transformer-based phonetic aligner

Last update: Dec 09, 2022

Related tags

Overview

Charsiu: A transformer-based phonetic aligner [arXiv]

Note. This is a preview version. The aligner is under active development. New functions, new languages and detailed documentation will be added soon!

Intro

Charsiu is a phonetic alignment tool, which can:

recognise phonemes in a given audio file
perform forced alignment using phone transcriptions created in the previous step or provided by the user.
directly predict the phone-to-audio alignment from audio (text-independent alignment)

Fun fact: Char Siu is one of the most representative dishes of Cantonese cuisine 🍲 (see wiki).

Tutorial (In progress)

You can directly run our model in the cloud via Google Colab!

Forced alignment:
Textless alignmnet:

Development plan

Package

Items	Progress
Documentation	Nov 2021
Textgrid support	Nov 2021
Model compression	TBD

Multilingual support

Language	Progress
English (American)	√
Mandarin Chinese	Nov 2021
Spanish	Dec 2021
English (British)	TBD
Cantonese	TBD
AAVE	TBD

Pretrained models

Our pretrained models are availble at the HuggingFace model hub: https://huggingface.co/charsiu.

Dependencies

pytorch
transformers
datasets
librosa
g2pe
praatio

Training

Coming soon!

Finetuning

Coming soon!

Attribution and Citation

For now, you can cite this tool as:

@article{zhu2019charsiu,
  title={Phone-to-audio alignment without text: A Semi-supervised Approach},
  author={Zhu, Jian and Zhang, Cong and Jurgens, David},
  journal={arXiv preprint arXiv:????????????????????},
  year={2021}
 }

To share a direct web link: https://github.com/lingjzhu/charsiu/.

References

Transformers
s3prl
Montreal Forced Aligner

Disclaimer

This tool is a beta version and is still under active development. It may have bugs and quirks, alongside the difficulties and provisos which are described throughout the documentation. This tool is distributed under MIT liscence. Please see license for details.

By using this tool, you acknowledge:

That you understand that this tool does not produce perfect camera-ready data, and that all results should be hand-checked for sanity's sake, or at the very least, noise should be taken into account.
That you understand that this tool is a work in progress which may contain bugs. Future versions will be released, and bug fixes (and additions) will not necessarily be advertised.
That this tool may break with future updates of the various dependencies, and that the authors are not required to repair the package when that happens.
That you understand that the authors are not required or necessarily available to fix bugs which are encountered (although you're welcome to submit bug reports to Jian Zhu ([email protected]), if needed), nor to modify the tool to your needs.
That you will acknowledge the authors of the tool if you use, modify, fork, or re-use the code in your future work.
That rather than re-distributing this tool to other researchers, you will instead advise them to download the latest version from the website.

... and, most importantly:

That neither the authors, our collaborators, nor the the University of Michigan or any related universities on the whole, are responsible for the results obtained from the proper or improper usage of the tool, and that the tool is provided as-is, as a service to our fellow linguists.

All that said, thanks for using our tool, and we hope it works wonderfully for you!

Support or Contact

Please contact Jian Zhu ([email protected]) for technical support.
Contact Cong Zhang ([email protected]) if you would like to receive more instructions on how to use the package.

Charsiu: A transformer-based phonetic aligner

Related tags

Overview

Charsiu: A transformer-based phonetic aligner [arXiv]

Intro

Tutorial (In progress)

Development plan

Pretrained models

Dependencies

Training

Finetuning

Attribution and Citation

References

Disclaimer

Support or Contact

Owner

jzhu

Python版OpenCVのTracking APIのサンプルです。DaSiamRPNアルゴリズムまで対応しています。

This repository contains the DendroMap implementation for scalable and interactive exploration of image datasets in machine learning.

Tensors and Dynamic neural networks in Python with strong GPU acceleration

Repo for paper "Dynamic Placement of Rapidly Deployable Mobile Sensor Robots Using Machine Learning and Expected Value of Information"

NLP made easy

A simple but complete full-attention transformer with a set of promising experimental features from various papers

PaddleViT: State-of-the-art Visual Transformer and MLP Models for PaddlePaddle 2.0+

Code for CPM-2 Pre-Train

Python package to generate image embeddings with CLIP without PyTorch/TensorFlow

TensorFlow 2 AI/ML library wrapper for openFrameworks

A U-Net combined with a variational auto-encoder that is able to learn conditional distributions over semantic segmentations.

ANEA: Distant Supervision for Low-Resource Named Entity Recognition

🐤 Nix-TTS: An Incredibly Lightweight End-to-End Text-to-Speech Model via Non End-to-End Distillation

(Py)TOD: Tensor-based Outlier Detection, A General GPU-Accelerated Framework

Elastic weight consolidation technique for incremental learning.

python library for invisible image watermark (blind image watermark)

You can draw the corresponding bounding box into the image and save it according to the result file (txt format) run by the tracker.

ISBI 2022: Cross-level Contrastive Learning and Consistency Constraint for Semi-supervised Medical Image.

Boundary-preserving Mask R-CNN (ECCV 2020)

Using contrastive learning and OpenAI's CLIP to find good embeddings for images with lossy transformations