Re-TACRED: Addressing Shortcomings of the TACRED Dataset

Last update: Dec 10, 2022

Related tags

Overview

Re-TACRED

Re-TACRED: Addressing Shortcomings of the TACRED Dataset
George Stoica, Emmanouil Antonios Platanios, and Barnabás Póczos
In Proceedings of the Thirty-fifth AAAI Conference on Artificial Intelligence 2021

Primary Contact: George Stoica. As of Jan 2021, I am no longer at CMU, and the cs.cmu.edu email may no longer work. Please contact me instead at: [email protected].

Changelog

1.0 - Initial dataset release: Data consisted of 105,206 total instances spread across 40 relations.
1.1 - Updated dataset release: After extensive discussion, we have elected to prune Re-TACRED by ~ 14K instances. The new dataset has 91,467 instances, spread across 40 relations. Pruned data consisted of a mixture of messily segmented entities (and corresponding types), or sentences whose relations were ambigious. While this version is smaller, it is cleaner, and better defined.

This repository contains all relevant resources for using Re-TACRED, a new relation extraction dataset.

For details on this work please check out our:

arXiv: Paper
AAAI 2021: Paper & Poster
NeurIPS 2020 KR2ML Workshop: Paper & Poster

Below we describe the contents of the four repository directories by name.

Re-TACRED

This directory contains version 1.1 of our revised TACRED dataset patches for each split. Due to licensing restrictions, we cannot provide the complete dataset. However, following Alt, Gabryszak, and Hennig (2020), our patch consists of json files mapping TACRED instances by their id to our revised labels.

The original TACRED dataset is available for download from the LDC here. It is free for members, or $25 for non-members.

Applying the patch is simple and only requires replacing each TACRED instance (where applicable) with our revised relation. For convenience, we provide a script for this named apply_patch.py in the Re-TACRED directory. In the script, you only need to replace

tacred_dir = None
save_dir = None

With the path to your TACRED dataset save directory, and the directory where you wish to save the patched data to respectively.

PA-LSTM, C-GCN & SpanBERT

We base our experiments off of the open-source model repositories of:

PA-LSTM: Zhang et. al. (2017)
C-GCN: Zhang et. al. (2018)
SpanBERT: Joshi et. al. (2019)

However, it is not possible to simply pass Re-TACRED to each model repository because each is hardcoded for TACRED. Thus, we must modify certain files to make each model Re-TACRED compatible. To make it as easy as possible, we provide all our altered files in each named model directory (e.g., the provided PA-LSTM directory). All that needs to be done is to replace the corresponding file in our provided directory with the corresponding file in the original model repository. For instance, you may replace SpanBERT's "run_tacred.py" file with our "run_tacred.py" file. Running experiments is equivalent to how it is performed in the original model repositories.

Note that our files also contain certain "quality of life" changes that make running each model more convenient for us. Examples include adding and tracking the test split while training (as opposed to only the dev set).

Re-TACRED: Addressing Shortcomings of the TACRED Dataset

Related tags

Overview

Re-TACRED

Owner

George Stoica

Face Mask Detection on Image and Video using tensorflow and keras

Numbering permanent and deciduous teeth via deep instance segmentation in panoramic X-rays

Highway networks implemented in PyTorch.

PyImpetus is a Markov Blanket based feature subset selection algorithm that considers features both separately and together as a group in order to provide not just the best set of features but also the best combination of features

Multi-Joint dynamics with Contact. A general purpose physics simulator.

A dead simple python wrapper for darknet that works with OpenCV 4.1, CUDA 10.1

The official repo of the CVPR2021 oral paper: Representative Batch Normalization with Feature Calibration

Tiny Object Detection in Aerial Images.

Machine Learning Privacy Meter: A tool to quantify the privacy risks of machine learning models with respect to inference attacks, notably membership inference attacks

Stable Neural ODE with Lyapunov-Stable Equilibrium Points for Defending Against Adversarial Attacks

A basic duplicate image detection service using perceptual image hash functions and nearest neighbor search, implemented using faiss, fastapi, and imagehash

EvDistill: Asynchronous Events to End-task Learning via Bidirectional Reconstruction-guided Cross-modal Knowledge Distillation (CVPR'21)

Telegram chatbot created with deep learning model (LSTM) and telebot library.

Learning where to learn - Gradient sparsity in meta and continual learning

Repo for "Physion: Evaluating Physical Prediction from Vision in Humans and Machines" submission to NeurIPS 2021 (Datasets & Benchmarks track)

Edge Restoration Quality Assessment

Official repository for Fourier model that can generate periodic signals

StyleSwin: Transformer-based GAN for High-resolution Image Generation

UnpNet - Rethinking 3-D LiDAR Point Cloud Segmentation(IEEE TNNLS)

Pca-on-genotypes - Mini bioinformatics project - PCA on genotypes