VilaQuAD: an extractive QA dataset from Catalan newswire
收藏资源简介:
<em>Dataset de QA extractiu amb 6282 parells de pregunta-resposta desenvolupats a partir de paràgrafs del diari en línia Vilaweb (https://www.vilaweb.cat).</em> From a the online edition of the catalan newspaper Vilaweb (https://www.vilaweb.cat), 2095 articles were randomnly selected. These headlines were also used to create a Textual Entailment dataset. For the extractive QA dataset, creation of between 1 and 5 questions for each news context was commissioned, following an adaptation of the guidelines from SQUAD 1.0 (Rajpurkar, Pranav et al. “SQuAD: 100, 000+ Questions for Machine Comprehension of Text.” EMNLP (2016)), http://arxiv.org/abs/1606.05250. In total, 6282 pairs of a question and an extracted fragment that contains the answer were created. Copyright (c) 2021 Text Mining Unit at BSC Funded by the Generalitat de Catalunya, Departament de Polítiques Digitals i Administració Pública (AINA), MT4ALL and Plan de Impulso de las Tecnologías del Lenguaje (Plan TL).



