allenai/openbookqa
收藏资源简介:
OpenBookQA数据集旨在促进高级问答研究,深入理解主题和语言表达。它包含需要多步推理、使用额外常识和丰富文本理解的问题。数据集分为main和additional两个配置,每个配置包含训练、验证和测试三个分割。数据字段包括问题ID、问题主干、选项、答案键等,additional配置还包含相关事实、人类评分、清晰度评分和匿名工作者ID。数据集的大小在1K到10K之间,语言为英语,任务类别为问答,任务ID为开放域问答。
The OpenBookQA dataset is designed to advance advanced question answering (QA) research and foster in-depth understanding of topics and linguistic expressions. It contains questions that require multi-step reasoning, external common sense utilization, and comprehensive text comprehension. The dataset is split into two configurations: main and additional, each consisting of three splits: training, validation, and test. Its data fields include question ID, question stem, options, answer key, etc. The additional configuration additionally includes supporting facts, human ratings, clarity scores, and anonymous worker IDs. The dataset has a size ranging from 1K to 10K, uses English as its language, belongs to the question answering task category, and its task ID is open-domain QA.
数据集概述
数据集名称: OpenBookQA
语言: 英语 (en)
许可证: 未知
多语言性: 单语
大小类别: 1K<n<10K
源数据集: 原始
任务类别: 问答
任务ID: open-domain-qa
论文代码ID: openbookqa
美观名称: OpenBookQA
数据集结构
数据实例
-
main配置:
id: 字符串类型question_stem: 字符串类型choices: 字典类型,包含text(字符串类型)和label(字符串类型)answerKey: 字符串类型
-
additional配置:
id: 字符串类型question_stem: 字符串类型choices: 字典类型,包含text(字符串类型)和label(字符串类型)answerKey: 字符串类型fact1: 字符串类型humanScore: 浮点数类型clarity: 浮点数类型turkIdAnonymized: 字符串类型
数据分割
| 名称 | 训练 | 验证 | 测试 |
|---|---|---|---|
| main | 4957 | 500 | 500 |
| additional | 4957 | 500 | 500 |
数据集创建
注释创建者:
- 众包
- 专家生成
语言创建者:
- 专家生成




