JaQuAD: Japanese Question Answering Dataset

Last update: Dec 27, 2022

Related tags

Overview

JaQuAD: Japanese Question Answering Dataset

Overview

Japanese Question Answering Dataset (JaQuAD), released in 2022, is a human-annotated dataset created for Japanese Machine Reading Comprehension. JaQuAD is developed to provide a SQuAD-like QA dataset in Japanese. JaQuAD contains 39,696 question-answer pairs. Questions and answers are manually curated by human annotators. Contexts are collected from Japanese Wikipedia articles.

For more information on how the dataset was created, refer to our paper, JaQuAD: Japanese Question Answering Dataset for Machine Reading Comprehension.

Data

JaQuAD consists of three sets: train, validation, and test. They were created from disjoint sets of Wikipedia articles. The following table shows statistics for each set:

Set	Number of Articles	Number of Contexts	Number of Questions
Train	691	9713	31748
Validation	101	1431	3939
Test	109	1479	4009

You can also download our dataset here. (The test set is not publicly released yet.)

from datasets import load_dataset
jaquad_data = load_dataset('SkelterLabsInc/JaQuAD')

Baseline

We also provide a baseline model for JaQuAD for comparison. We created this model by fine-tuning a publicly available Japanese BERT model on JaQuAD. You can see the performance of the baseline model in the table below.

For more information on the model's creation, refer to JaQuAD.ipynb.

Pre-trained LM	Dev F1	Dev EM	Test F1	Test EM
BERT-Japanese	77.35	61.01	78.92	63.38

You can download the baseline model here.

Usage

from transformers import AutoModelForQuestionAnswering, AutoTokenizer

question = 'アレクサンダー・グラハム・ベルは、どこで生まれたの?'
context = 'アレクサンダー・グラハム・ベルは、スコットランド生まれの科学者、発明家、工学者である。世界初の>実用的電話の発明で知られている。'

model = AutoModelForQuestionAnswering.from_pretrained(
    'SkelterLabsInc/bert-base-japanese-jaquad')
tokenizer = AutoTokenizer.from_pretrained(
    'SkelterLabsInc/bert-base-japanese-jaquad')

inputs = tokenizer(
    question, context, add_special_tokens=True, return_tensors="pt")
input_ids = inputs["input_ids"].tolist()[0]
outputs = model(**inputs)
answer_start_scores = outputs.start_logits
answer_end_scores = outputs.end_logits

# Get the most likely start of the answer with the argmax of the score.
answer_start = torch.argmax(answer_start_scores)
# Get the most likely end of the answer with the argmax of the score.
# 1 is added to `answer_end` because the index of the score is inclusive.
answer_end = torch.argmax(answer_end_scores) + 1

answer = tokenizer.convert_tokens_to_string(
    tokenizer.convert_ids_to_tokens(input_ids[answer_start:answer_end]))
# answer = 'スコットランド'

Limitations

This dataset is not yet complete. The social biases of this dataset have not yet been investigated.

If you find any errors in JaQuAD, please contact [email protected].

Reference

If you use our dataset or code, please cite our paper:

@misc{so2022jaquad,
      title={{JaQuAD: Japanese Question Answering Dataset for Machine Reading Comprehension}},
      author={ByungHoon So and Kyuhong Byun and Kyungwon Kang and Seongjin Cho},
      year={2022},
      eprint={2202.01764},
      archivePrefix={arXiv},
      primaryClass={cs.CL}
}

LICENSE

The JaQuAD dataset is licensed under the [CC BY-SA 3.0] (https://creativecommons.org/licenses/by-sa/3.0/) license.

Have Questions?

Ask us at [email protected].

JaQuAD: Japanese Question Answering Dataset

Related tags

Overview

JaQuAD: Japanese Question Answering Dataset

Overview

Data

Baseline

Usage

Limitations

Reference

LICENSE

Have Questions?

Owner

SkelterLabs

Pre-training with Extracted Gap-sentences for Abstractive SUmmarization Sequence-to-sequence models

A curated list of FOSS tools to improve the Hacker News experience

Creating an Audiobook (mp3 file) using a Ebook (epub) using BeautifulSoup and Google Text to Speech

(ACL 2022) The source code for the paper "Towards Abstractive Grounded Summarization of Podcast Transcripts"

A python script to prefab your scripts/text files, and re create them with ease and not have to open your browser to copy code or write code yourself

SDL: Synthetic Document Layout dataset

Fuzzy String Matching in Python

voice2json is a collection of command-line tools for offline speech/intent recognition on Linux

Pytorch version of BERT-whitening

Multilingual word vectors in 78 languages

A fast Text-to-Speech (TTS) model. Work well for English, Mandarin/Chinese, Japanese, Korean, Russian and Tibetan (so far). 快速语音合成模型，适用于英语、普通话/中文、日语、韩语、俄语和藏语（当前已测试）。

Amazon Multilingual Counterfactual Dataset (AMCD)

An open source library for deep learning end-to-end dialog systems and chatbots.

The following links explain a bit the idea of semantic search and how search mechanisms work by doing retrieve and rerank

Unofficial Python library for using the Polish Wordnet (plWordNet / Słowosieć)

A modular Karton Framework service that unpacks common packers like UPX and others using the Qiling Framework.

DeepPavlov Tutorials

Data and code to support "Applied Natural Language Processing" (INFO 256, Fall 2021, UC Berkeley)

CMeEE 数据集医学实体抽取

TPlinker for NER 中文/英文命名实体识别