OpenBookQA
收藏资源简介:
“OpenBookQA 是一种新的问答数据集,它以开卷考试为模型,用于评估人类对学科的理解。它由 5,957 个多项选择的初级科学问题(4,957 个训练,500 个开发,500 个测试)组成,它探讨了对 1,326 个核心科学事实的小“书”的理解以及这些事实在新情况中的应用。对于训练,数据集包括从每个问题到它旨在探索的核心科学事实的映射。回答 OpenBookQA 问题需要书中未包含的其他广泛的常识。这些问题在设计上会被基于检索的算法和单词共现算法错误地回答。此外,数据集包括 5,167 个众包常识的集合事实,以及训练/开发/测试问题的扩展版本,其中每个问题都与其原始核心事实、人类准确性分数、清晰度分数和匿名人群相关联rker ID。”
OpenBookQA is a novel question answering dataset modeled after open-book exams, designed to evaluate human subject matter understanding. It comprises 5,957 multiple-choice middle school science questions (4,957 for training, 500 for development, and 500 for test), which probe the understanding of a small "book" of 1,326 core scientific facts and the application of these facts in novel scenarios. For the training split, the dataset includes mappings from each question to the core scientific facts that it aims to explore. Answering OpenBookQA questions requires additional broad common knowledge not contained within the "book". These questions are intentionally designed to be incorrectly answered by retrieval-based algorithms and word co-occurrence based models. Additionally, the dataset includes a collection of 5,167 crowdsourced common knowledge facts, as well as extended versions of the train/dev/test questions, where each question is linked to its original core fact, human accuracy score, clarity score, and anonymized worker ID.

- OpenBookQA数据集首次发表,由美国艾伦人工智能研究所(Allen Institute for AI)发布,旨在评估机器在开放领域问答中的理解能力。
- OpenBookQA数据集首次应用于机器学习竞赛,吸引了全球多个研究团队参与,推动了问答系统技术的发展。
- OpenBookQA数据集被广泛应用于学术研究,成为评估自然语言处理模型性能的重要基准之一。
- OpenBookQA数据集的扩展版本发布,增加了更多复杂问题和多样的知识领域,进一步提升了数据集的挑战性和实用性。



