EstQA Question Answering dataset
收藏资源简介:
Dataset for extractive question answering in Estonian. It based on Wikipedia articles, pre-filtered via PageRank. Training set includes 776 context-question-answer triplets. There are several possible answers per question, each in a separate triplet. Number of different questions is 512. Test set includes 603 samples. Each sample contains one or more golden answers. Altogether there are 892 golden answers. If you use this dataset for research, please cite the following paper: @mastersthesis{mastersthesis, author = {Anu Käver}, title = {Extractive Question Answering for Estonian Language}, school = {Tallinn University of Technology (TalTech)}, year = 2021 }
本数据集为爱沙尼亚语抽取式问答(Extractive Question Answering)数据集,其基于维基百科(Wikipedia)文章构建,并通过PageRank算法完成预筛选。 训练集共包含776组上下文-问题-答案三元组。每个问题对应多个可选答案,且每个答案单独作为一条三元组;训练集中总计包含512个不同的问题。 测试集共包含603个样本,每个样本包含一个或多个标准答案,所有样本累计共有892条标准答案。 若将本数据集用于研究工作,请引用如下学位论文: @mastersthesis{mastersthesis, author = {Anu Käver}, title = {面向爱沙尼亚语的抽取式问答(Extractive Question Answering for Estonian Language)}, school = {塔林理工大学(Tallinn University of Technology, TalTech)}, year = 2021 }



