QAngaroo/wiki_hop
收藏资源简介:
--- annotations_creators: - crowdsourced language_creators: - expert-generated language: - en license: - cc-by-sa-3.0 multilinguality: - monolingual size_categories: - 10K<n<100K source_datasets: - original task_categories: - question-answering task_ids: - extractive-qa paperswithcode_id: wikihop pretty_name: WikiHop tags: - multi-hop dataset_info: - config_name: original features: - name: id dtype: string - name: query dtype: string - name: answer dtype: string - name: candidates sequence: string - name: supports sequence: string - name: annotations sequence: sequence: string splits: - name: train num_bytes: 325952974 num_examples: 43738 - name: validation num_bytes: 41246536 num_examples: 5129 download_size: 339843061 dataset_size: 367199510 - config_name: masked features: - name: id dtype: string - name: question dtype: string - name: answer dtype: string - name: candidates sequence: string - name: supports sequence: string - name: annotations sequence: sequence: string splits: - name: train num_bytes: 348249138 num_examples: 43738 - name: validation num_bytes: 44066862 num_examples: 5129 download_size: 339843061 dataset_size: 392316000 --- # Dataset Card for WikiHop ## Table of Contents - [Dataset Description](#dataset-description) - [Dataset Summary](#dataset-summary) - [Supported Tasks and Leaderboards](#supported-tasks-and-leaderboards) - [Languages](#languages) - [Dataset Structure](#dataset-structure) - [Data Instances](#data-instances) - [Data Fields](#data-fields) - [Data Splits](#data-splits) - [Dataset Creation](#dataset-creation) - [Curation Rationale](#curation-rationale) - [Source Data](#source-data) - [Annotations](#annotations) - [Personal and Sensitive Information](#personal-and-sensitive-information) - [Considerations for Using the Data](#considerations-for-using-the-data) - [Social Impact of Dataset](#social-impact-of-dataset) - [Discussion of Biases](#discussion-of-biases) - [Other Known Limitations](#other-known-limitations) - [Additional Information](#additional-information) - [Dataset Curators](#dataset-curators) - [Licensing Information](#licensing-information) - [Citation Information](#citation-information) - [Contributions](#contributions) ## Dataset Description - **Homepage:** [QAngaroo](http://qangaroo.cs.ucl.ac.uk/) - **Repository:** [If the dataset is hosted on github or has a github homepage, add URL here]() - **Paper:** [Constructing Datasets for Multi-hop Reading Comprehension Across Documents](https://arxiv.org/abs/1710.06481) - **Leaderboard:** [leaderboard](http://qangaroo.cs.ucl.ac.uk/leaderboard.html) - **Point of Contact:** [Johannes Welbl](j.welbl@cs.ucl.ac.uk) ### Dataset Summary [More Information Needed] ### Supported Tasks and Leaderboards [More Information Needed] ### Languages [More Information Needed] ## Dataset Structure ### Data Instances [More Information Needed] ### Data Fields [More Information Needed] ### Data Splits [More Information Needed] ## Dataset Creation ### Curation Rationale [More Information Needed] ### Source Data [More Information Needed] #### Initial Data Collection and Normalization [More Information Needed] #### Who are the source language producers? [More Information Needed] ### Annotations [More Information Needed] #### Annotation process [More Information Needed] #### Who are the annotators? [More Information Needed] ### Personal and Sensitive Information [More Information Needed] ## Considerations for Using the Data ### Social Impact of Dataset [More Information Needed] ### Discussion of Biases [More Information Needed] ### Other Known Limitations [More Information Needed] ## Additional Information ### Dataset Curators [More Information Needed] ### Licensing Information [More Information Needed] ### Citation Information [More Information Needed] ### Contributions Thanks to [@patil-suraj](https://github.com/patil-suraj) for adding this dataset.
annotations_creators: - 众包(crowdsourced) language_creators: - 专家生成(expert-generated) language: - 英语 license: - CC BY-SA 3.0 multilinguality: - 单语言 size_categories: - 10K<n<100K source_datasets: - 原创 task_categories: - 问答(question-answering) task_ids: - 抽取式问答(extractive-qa) paperswithcode_id: wikihop pretty_name: WikiHop tags: - 多跳(multi-hop) dataset_info: - config_name: 原版(original) features: - name: id dtype: 字符串(string) - name: query dtype: 字符串(string) - name: answer dtype: 字符串(string) - name: candidates dtype: 字符串序列 - name: supports dtype: 字符串序列 - name: annotations dtype: 二维字符串序列 splits: - name: 训练集(train) num_bytes: 325952974 num_examples: 43738 - name: 验证集(validation) num_bytes: 41246536 num_examples: 5129 download_size: 339843061 dataset_size: 367199510 - config_name: 掩码版(masked) features: - name: id dtype: 字符串(string) - name: question dtype: 字符串(string) - name: answer dtype: 字符串(string) - name: candidates dtype: 字符串序列 - name: supports dtype: 字符串序列 - name: annotations dtype: 二维字符串序列 splits: - name: 训练集(train) num_bytes: 348249138 num_examples: 43738 - name: 验证集(validation) num_bytes: 44066862 num_examples: 5129 download_size: 339843061 dataset_size: 392316000 --- # WikiHop 数据集卡片 ## 目录 - [数据集描述](#dataset-description) - [数据集概述](#dataset-summary) - [支持任务与排行榜](#supported-tasks-and-leaderboards) - [语言](#languages) - [数据集结构](#dataset-structure) - [数据实例](#data-instances) - [数据字段](#data-fields) - [数据拆分](#data-splits) - [数据集构建](#dataset-creation) - [构建初衷](#curation-rationale) - [源数据](#source-data) - [标注信息](#annotations) - [个人与敏感信息](#personal-and-sensitive-information) - [数据使用注意事项](#considerations-for-using-the-data) - [数据集的社会影响](#social-impact-of-dataset) - [偏差讨论](#discussion-of-biases) - [其他已知局限](#other-known-limitations) - [附加信息](#additional-information) - [数据集维护者](#dataset-curators) - [许可信息](#licensing-information) - [引用信息](#citation-information) - [贡献](#contributions) ## 数据集描述 - **主页:** [QAngaroo](http://qangaroo.cs.ucl.ac.uk/) - **代码仓库:** [若数据集托管于GitHub或拥有GitHub主页,请在此添加链接]() - **论文:** [跨文档多跳阅读理解数据集构建](https://arxiv.org/abs/1710.06481) - **排行榜:** [排行榜](http://qangaroo.cs.ucl.ac.uk/leaderboard.html) - **联络人:** [约翰内斯·韦尔布(Johannes Welbl)](j.welbl@cs.ucl.ac.uk) ### 数据集概述 [需补充更多信息] ### 支持任务与排行榜 [需补充更多信息] ### 语言 [需补充更多信息] ## 数据集结构 ### 数据实例 [需补充更多信息] ### 数据字段 [需补充更多信息] ### 数据拆分 [需补充更多信息] ## 数据集构建 ### 构建初衷 [需补充更多信息] ### 源数据 [需补充更多信息] #### 初始数据收集与标准化 [需补充更多信息] #### 源语言生产者是谁? [需补充更多信息] ### 标注信息 [需补充更多信息] #### 标注流程 [需补充更多信息] #### 标注者是谁? [需补充更多信息] ### 个人与敏感信息 [需补充更多信息] ## 数据使用注意事项 ### 数据集的社会影响 [需补充更多信息] ### 偏差讨论 [需补充更多信息] ### 其他已知局限 [需补充更多信息] ## 附加信息 ### 数据集维护者 [需补充更多信息] ### 许可信息 [需补充更多信息] ### 引用信息 [需补充更多信息] ### 贡献 感谢 [@patil-suraj](https://github.com/patil-suraj) 贡献此数据集。
数据集卡片 for WikiHop
数据集描述
- annotations_creators: crowdsourced
- language_creators: expert-generated
- language: en
- license: cc-by-sa-3.0
- multilinguality: monolingual
- size_categories: 10K<n<100K
- source_datasets: original
- task_categories: question-answering
- task_ids: extractive-qa
- paperswithcode_id: wikihop
- pretty_name: WikiHop
- tags: multi-hop
数据集结构
配置信息
原始配置
- features:
- id: string
- query: string
- answer: string
- candidates: sequence: string
- supports: sequence: string
- annotations: sequence: sequence: string
- splits:
- train:
- num_bytes: 325952974
- num_examples: 43738
- validation:
- num_bytes: 41246536
- num_examples: 5129
- train:
- download_size: 339843061
- dataset_size: 367199510
掩码配置
- features:
- id: string
- question: string
- answer: string
- candidates: sequence: string
- supports: sequence: string
- annotations: sequence: sequence: string
- splits:
- train:
- num_bytes: 348249138
- num_examples: 43738
- validation:
- num_bytes: 44066862
- num_examples: 5129
- train:
- download_size: 339843061
- dataset_size: 392316000




