HybridQA
收藏资源简介:
HybridQA是由加州大学圣巴巴拉分校创建的一个大规模多跳问答数据集,旨在解决现有数据集在处理异构信息时的覆盖问题。该数据集包含约70,000个问答对,每个问题都与一个维基百科表格和多个与表格实体相关的自由形式文本相关联。数据集的设计要求模型在回答问题时必须结合表格信息和文本信息,这使得HybridQA成为一个挑战性的基准,用于研究在异构信息上的问答系统。
HybridQA is a large-scale multi-hop question answering dataset created by the University of California, Santa Barbara, aiming to address the coverage limitation of existing datasets when handling heterogeneous information. This dataset contains approximately 70,000 question-answer pairs, where each question is associated with a Wikipedia table and multiple free-form texts related to the entities in the table. The dataset is designed such that models must integrate both tabular and textual information to answer the questions, making HybridQA a challenging benchmark for research on question answering systems over heterogeneous information.

- 1HybridQA: A Dataset of Multi-Hop Question Answering over Tabular and Textual Data加州大学圣巴巴拉分校 · 2021年



