Use PaddlePaddle to reproduce the paper：mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer

Last update: Oct 17, 2021

Related tags

Text Data & NLP MT5_paddle

Overview

MT5_paddle

Use PaddlePaddle to reproduce the paper：mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer

English | 简体中文

mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer

Abstract： The recent “Text-to-Text Transfer Transformer” (T5) leveraged a unified text-to-text format and scale to attain state-of-the-art results on a wide variety of English-language NLP tasks. In this paper, we introduce mT5, a multilingual variant of T5 that was pre-trained on a new Common Crawl-based dataset covering 101 languages. We detail the design and modified training of mT5 and demonstrate its state-of-the-art performance on many multilingual benchmarks. We also describe a simple technique to prevent “accidental translation” in the zero-shot setting, where a generative model chooses to (partially) translate its prediction into the wrong language. All of the code and model checkpoints used in this work are publicly available.

This project is an open source implementation of MT5 on Paddle 2.x.

Environment Installation

label	value
python	>=3.6
GPU	V100
Frame	PaddlePaddle2.1.2
Cuda	10.1
Cudnn	7.6

Cloud platform used in this recurrence：https://aistudio.baidu.com/

# Clone the repository
git clone https://github.com/27182812/MT5_paddle
# Enter the root directory
cd MT5_paddle
# Install the necessary python libraries locally
pip install -r requirements.txt

"test.ipynb" has run results display.

Quick Start

（一）Tokenizer Accuracy Alignment

### 对齐tokenizer
text = "Welcome to use paddle and paddlenlp!"
torch_tokenizer = PTT5Tokenizer.from_pretrained("./mt5-large")
paddle_tokenizer = PDT5Tokenizer.from_pretrained("./mt5-large")
torch_inputs = torch_tokenizer(text)
paddle_inputs = paddle_tokenizer(text)
print(torch_inputs)
print(paddle_inputs)

（二）Model Accuracy Alignment

run python compare.py，Comparing the accuracy between huggingface and paddle.

python compare.py
# MT5-large-pytorch vs paddle MT5-large-paddle
mean difference: tensor(2.0390e-06)
max difference: tensor(0.0004)

(三）Weights Transform

run python convert.py，transform weights of huggingface model to weights of paddle model. The weight path needs to be replaced

(四）Downstream task fine-tuning

run python train.py. "args.py" is for parameter.

Reference

大佬的T5代码：https://github.com/JunnYu/paddle_t5

@unknown{unknown,
author = {Xue, Linting and Constant, Noah and Roberts, Adam and Kale, Mihir and Al-Rfou, Rami and Siddhant, Aditya and Barua, Aditya and Raffel, Colin},
year = {2020},
month = {10},
pages = {},
title = {mT5: A massively multilingual pre-trained text-to-text transformer}
}

Use PaddlePaddle to reproduce the paper：mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer

Related tags

Overview

MT5_paddle

Environment Installation

Quick Start

（一）Tokenizer Accuracy Alignment

（二）Model Accuracy Alignment

(三）Weights Transform

(四）Downstream task fine-tuning

Reference

Owner

Basic yet complete Machine Learning pipeline for NLP tasks

A fast Text-to-Speech (TTS) model. Work well for English, Mandarin/Chinese, Japanese, Korean, Russian and Tibetan (so far). 快速语音合成模型，适用于英语、普通话/中文、日语、韩语、俄语和藏语（当前已测试）。

AllenNLP integration for Shiba: Japanese CANINE model

A natural language modeling framework based on PyTorch

customer care chatbot made with Rasa Open Source.

BERT, LDA, and TFIDF based keyword extraction in Python

UniSpeech - Large Scale Self-Supervised Learning for Speech

Code for paper Multitask-Finetuning of Zero-shot Vision-Language Models

NAACL 2022: MCSE: Multimodal Contrastive Learning of Sentence Embeddings

IEEEXtreme15.0 Questions And Answers

SimCTG - A Contrastive Framework for Neural Text Generation

Extract Keywords from sentence or Replace keywords in sentences.

Use the power of GPT3 to execute any function inside your programs just by giving some doctests

👄 The most accurate natural language detection library for Python, suitable for long and short text alike

Unofficial Python library for using the Polish Wordnet (plWordNet / Słowosieć)

The entmax mapping and its loss, a family of sparse softmax alternatives.

Tools and data for measuring the popularity & growth of various programming languages.

PyTorch implementation of the paper: Text is no more Enough! A Benchmark for Profile-based Spoken Language Understanding

File-based TF-IDF: Calculates keywords in a document, using a word corpus.

Anomaly Detection 이상치 탐지 전처리 모듈