MultiLexNorm 2021 competition system from ÚFAL

Last update: Jun 28, 2022

Overview

ÚFAL at MultiLexNorm 2021:
Improving Multilingual Lexical Normalization by Fine-tuning ByT5

David Samuel & Milan Straka

Charles University
Faculty of Mathematics and Physics
Institute of Formal and Applied Linguistics

Paper (TODO)
Interactive demo on Google Colab
HuggingFace models (TODO)

This is the official repository for the winning entry to the W-NUT 2021: Multilingual Lexical Normalization (MultiLexNorm) shared task, which evaluates lexical-normalization systems on 12 social media datasets in 11 languages.

Our system is based on ByT5, which we first pre-train on synthetic data and then fine-tune on authentic normalization data. It achieves the best performance by a wide margin in intrinsic evaluation, and also the best performance in extrinsic evaluation through dependency parsing. In addition to these source files, we also release the fine-tuned models on HuggingFace (TODO) and an interactive demo on Google Colab.

How to run

🐾 Clone repository and install the Python requirements

git clone https://github.com/ufal/multilexnorm2021.git
cd multilexnorm2021

pip3 install -r requirements.txt

🐾 Initialize

Run the inialization script to download the official MultiLexNorm data together with a dump of English Wikipedia. We recommend downloading Wikipidia dumps to get clean multi-lingual data, but other data sources should also work.

./initialize.sh

🐾 Train

To train a model for English lexical normalization, simply run the following script. Other configurations are located in the config folder.

python3 train.py --config config/en.yaml

Please cite the following publication

@inproceedings{wnut-ufal,
  title= "{ÚFAL} at {MultiLexNorm} 2021: Improving Multilingual Lexical Normalization by Fine-tuning {ByT5}",
  author = "Samuel, David and Straka, Milan",
  booktitle = "Proceedings of the 7th Workshop on Noisy User-generated Text (W-NUT 2021)",
  year = "2021",
  publisher = "Association for Computational Linguistics",
  address = "Punta Cana, Dominican Republic"
}

You might also like...

My published benchmark for a Kaggle Simulations Competition

Lux AI Working Title Bot Please refer to the Kaggle notebook for the comment section. The comment section contains my explanation on my code structure

29 Aug 22, 2022

Top #1 Submission code for the first https://alphamev.ai MEV competition with best AUC (0.9893) and MSE (0.0982).

alphamev-winning-submission Top #1 Submission code for the first alphamev MEV competition with best AUC (0.9893) and MSE (0.0982). The code won't run

70 Oct 29, 2022

Omnidirectional Scene Text Detection with Sequential-free Box Discretization (IJCAI 2019). Including competition model, online demo, etc.

Box_Discretization_Network This repository is built on the pytorch [maskrcnn_benchmark]. The method is the foundation of our ReCTs-competition method

266 Nov 24, 2022

Team nan solution repository for FPT data-centric competition. Data augmentation, Albumentation, Mosaic, Visualization, KNN application

FPT_data_centric_competition - Team nan solution repository for FPT data-centric competition. Data augmentation, Albumentation, Mosaic, Visualization, KNN application

2 Oct 30, 2022

MultiLexNorm 2021 competition system from ÚFAL

Related tags

Overview

ÚFAL at MultiLexNorm 2021:Improving Multilingual Lexical Normalization by Fine-tuning ByT5

ÚFAL at MultiLexNorm 2021:

How to run

🐾 Clone repository and install the Python requirements

🐾 Initialize

🐾 Train

Please cite the following publication

You might also like...

My published benchmark for a Kaggle Simulations Competition

Top #1 Submission code for the first https://alphamev.ai MEV competition with best AUC (0.9893) and MSE (0.0982).

Omnidirectional Scene Text Detection with Sequential-free Box Discretization (IJCAI 2019). Including competition model, online demo, etc.

Team nan solution repository for FPT data-centric competition. Data augmentation, Albumentation, Mosaic, Visualization, KNN application

Solution of Kaggle competition: Sartorius - Cell Instance Segmentation

Job-Recommend-Competition - Vectorwise Interpretable Attentions for Multimodal Tabular Data

Group project for MFIN7036. Our goal is to predict firm profitability with text-based competition measures.

Data visualization app for H&M competition in kaggle

This is the solution for 2nd rank in Kaggle competition: Feedback Prize - Evaluating Student Writing.

Releases(v1.0.0)

v1.0.0(Dec 5, 2021)

Owner

ÚFAL

Pacman-AI - AI project designed by UC Berkeley. Designed reflex and minimax agents for the game Pacman.

Wordle Env: A Daily Word Environment for Reinforcement Learning

Implement Decoupled Neural Interfaces using Synthetic Gradients in Pytorch

Named Entity Recognition with Small Strongly Labeled and Large Weakly Labeled Data

Research on Event Accumulator Settings for Event-Based SLAM

This repo contains the source code and a benchmark for predicting user's utilities with Machine Learning techniques for Computational Persuasion

Source code for CVPR2022 paper "Abandoning the Bayer-Filter to See in the Dark"

An investigation project for SISR.

Paper Code：A Self-adaptive Weighted Differential Evolution Approach for Large-scale Feature Selection

Cl datasets - PyTorch image dataloaders and utility functions to load datasets for supervised continual learning

PyTorch Implementation of VAENAR-TTS: Variational Auto-Encoder based Non-AutoRegressive Text-to-Speech Synthesis.

pyspark🍒🥭 is delicious，just eat it!😋😋

The description of FMFCC-A (audio track of FMFCC) dataset and Challenge resluts.

MIMIC Code Repository: Code shared by the research community for the MIMIC-III database

Revisiting Discriminator in GAN Compression: A Generator-discriminator Cooperative Compression Scheme (NeurIPS2021)

The code for our paper "NSP-BERT: A Prompt-based Zero-Shot Learner Through an Original Pre-training Task —— Next Sentence Prediction"

Baleen: Robust Multi-Hop Reasoning at Scale via Condensed Retrieval (NeurIPS'21)

Company clustering with K-means/GMM and visualization with PCA, t-SNE, using SSAN relation extraction

Semi-supervised Adversarial Learning to Generate Photorealistic Face Images of New Identities from 3D Morphable Model

PyTorch implementation of InstaGAN: Instance-aware Image-to-Image Translation

ÚFAL at MultiLexNorm 2021:
Improving Multilingual Lexical Normalization by Fine-tuning ByT5