DuReader
收藏资源简介:
DuReader是由百度公司创建的一个大规模、开放领域的中国机器阅读理解(MRC)数据集,旨在解决实际的MRC问题。该数据集包含20万问题、42万答案和100万文档,是目前最大的中文MRC数据集。数据来源于百度搜索和百度知道,答案由人工生成,特别强调了yes-no和意见问题。DuReader的创建过程涉及从搜索日志中随机抽样问题,并使用预训练分类器自动选择问题,然后由人工标注。该数据集主要应用于机器阅读理解领域,旨在推动MRC技术的发展,特别是在处理真实世界数据源、多样问题类型和大规模数据方面。
DuReader is a large-scale, open-domain Chinese machine reading comprehension (MRC) dataset developed by Baidu, which is designed to address real-world MRC problems. It contains 200,000 questions, 420,000 answers and 1,000,000 documents, making it the largest Chinese MRC dataset to date. The dataset is sourced from Baidu Search and Baidu Zhidao, with all answers manually generated, and special emphasis is placed on yes-no questions and opinion-based questions. The creation process of DuReader includes randomly sampling questions from search logs, automatically selecting the sampled questions via pre-trained classifiers, followed by manual annotation. This dataset is primarily applied in the field of machine reading comprehension, aiming to promote the development of MRC technologies, especially in handling real-world data sources, diverse question types and large-scale datasets.




