遇见数据集

VilaQuAD: an extractive QA dataset from Catalan newswire

收藏
Zenodo2021-02-25 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

<em>Dataset de QA extractiu amb 6282 parells de pregunta-resposta desenvolupats a partir de paràgrafs del diari en línia Vilaweb (https://www.vilaweb.cat).</em> From a the online edition of the catalan newspaper Vilaweb (https://www.vilaweb.cat), 2095 articles were randomnly selected. These headlines were also used to create a Textual Entailment dataset. For the extractive QA dataset, creation of between 1 and 5 questions for each news context was commissioned, following an adaptation of the guidelines from SQUAD 1.0 (Rajpurkar, Pranav et al. “SQuAD: 100, 000+ Questions for Machine Comprehension of Text.” EMNLP (2016)), http://arxiv.org/abs/1606.05250. In total, 6282 pairs of a question and an extracted fragment that contains the answer were created. Copyright (c) 2021 Text Mining Unit at BSC Funded by the Generalitat de Catalunya, Departament de Polítiques Digitals i Administració Pública (AINA), MT4ALL and Plan de Impulso de las Tecnologías del Lenguaje (Plan TL).

提供机构:
Zenodo
创建时间:
2021-02-25
二维码
社区交流群
二维码
科研交流群
商业服务