DROP
收藏资源简介:
DROP数据集是由艾伦人工智能研究所创建的一个英语阅读理解基准,旨在推动对段落内容的更全面分析。该数据集包含96,567个问题,要求系统对段落内容进行离散推理,如加法、计数或排序等。这些问题需要比以往数据集更深入的段落理解。数据集通过众包创建,首先从维基百科收集易于提问的段落,然后鼓励众工作者提出挑战性问题。DROP数据集特别强调体育比赛摘要和历史文章,旨在推动结合分布式表示与符号离散推理的研究,解决阅读理解系统在复杂问题处理上的不足。
The DROP dataset is an English reading comprehension benchmark developed by the Allen Institute for AI, aiming to advance more comprehensive analysis of passage-level content. It contains 96,567 questions that require systems to perform discrete reasoning over passage content, such as addition, counting, sorting, and other similar operations. These questions demand deeper passage comprehension than those found in prior reading comprehension datasets. The dataset was built through crowdsourcing: first, passages suitable for question generation were collected from Wikipedia, and then crowdworkers were encouraged to develop challenging questions. The DROP dataset places particular emphasis on sports game recaps and historical articles, and is designed to advance research that combines distributed representations and symbolic discrete reasoning, addressing the limitations of reading comprehension systems in handling complex questions.

- DROP数据集首次发表,由华盛顿大学、艾伦人工智能研究所和卡内基梅隆大学共同发布。该数据集旨在推动阅读理解任务的发展,特别是针对需要复杂推理和计算的问答任务。
- DROP数据集在多个国际自然语言处理会议上被广泛讨论和应用,成为评估模型在复杂问答任务中表现的重要基准。
- 随着深度学习模型的进步,DROP数据集的应用范围进一步扩大,多个研究团队在其基础上提出了新的模型和方法,显著提升了阅读理解任务的性能。
- 1DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over ParagraphsUniversity of Washington, Allen Institute for AI · 2019年
- 2Improving Multi-hop Question Answering over Knowledge Graphs using Knowledge Base EmbeddingsUniversity of Waterloo, University of Toronto · 2020年
- 3UnifiedQA: Crossing Format Boundaries With a Single QA SystemAllen Institute for AI · 2020年
- 4A Simple and Effective Model for Answering Multi-span QuestionsUniversity of Washington, Allen Institute for AI · 2020年
- 5Multi-hop Question Answering via Reasoning ChainsUniversity of Illinois at Urbana-Champaign, Google Research · 2021年



