HOTPOTQA
收藏资源简介:
HOTPOTQA是一个包含113,000个基于维基百科的问题-答案对的大型数据集,由卡内基梅隆大学等机构创建。该数据集的特点在于其问题需要通过多文档推理来回答,且问题类型多样,不依赖于预先存在的知识库或知识模式。此外,HOTPOTQA提供了句子级别的支持事实,以帮助QA系统进行强监督推理和解释预测。数据集还引入了新的事实比较问题类型,以测试QA系统提取相关事实和进行必要比较的能力。HOTPOTQA的应用领域主要集中在测试和提升智能系统在自然语言处理中的多跳推理能力,旨在解决现有QA数据集在复杂推理和解释性方面的不足。
HOTPOTQA is a large-scale dataset consisting of 113,000 Wikipedia-based question-answer pairs, created by institutions including Carnegie Mellon University. The core feature of this dataset is that its questions require multi-document reasoning to answer, cover diverse question types, and do not rely on pre-existing knowledge bases or predefined knowledge schemas. In addition, HOTPOTQA provides sentence-level supporting facts to help QA systems conduct strongly supervised reasoning and interpretable prediction. Furthermore, the dataset introduces a novel factual comparison question type to evaluate the ability of QA systems to extract relevant factual information and perform necessary comparative reasoning. The main application scenarios of HOTPOTQA focus on testing and enhancing the multi-hop reasoning capability of intelligent systems in natural language processing, aiming to address the shortcomings of existing QA datasets in complex reasoning and explainability.

- 1HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering卡内基梅隆大学 · 2018年



