QAngaroo/med_hop
收藏资源简介:
--- annotations_creators: - crowdsourced language_creators: - expert-generated language: - en license: - cc-by-sa-3.0 multilinguality: - monolingual size_categories: - 1K<n<10K source_datasets: - original task_categories: - question-answering task_ids: - extractive-qa paperswithcode_id: medhop pretty_name: MedHop tags: - multi-hop dataset_info: - config_name: original features: - name: id dtype: string - name: query dtype: string - name: answer dtype: string - name: candidates sequence: string - name: supports sequence: string splits: - name: train num_bytes: 93937322 num_examples: 1620 - name: validation num_bytes: 16461640 num_examples: 342 download_size: 339843061 dataset_size: 110398962 - config_name: masked features: - name: id dtype: string - name: question dtype: string - name: answer dtype: string - name: candidates sequence: string - name: supports sequence: string splits: - name: train num_bytes: 95813584 num_examples: 1620 - name: validation num_bytes: 16800570 num_examples: 342 download_size: 339843061 dataset_size: 112614154 --- # Dataset Card for MedHop ## Table of Contents - [Dataset Description](#dataset-description) - [Dataset Summary](#dataset-summary) - [Supported Tasks and Leaderboards](#supported-tasks-and-leaderboards) - [Languages](#languages) - [Dataset Structure](#dataset-structure) - [Data Instances](#data-instances) - [Data Fields](#data-fields) - [Data Splits](#data-splits) - [Dataset Creation](#dataset-creation) - [Curation Rationale](#curation-rationale) - [Source Data](#source-data) - [Annotations](#annotations) - [Personal and Sensitive Information](#personal-and-sensitive-information) - [Considerations for Using the Data](#considerations-for-using-the-data) - [Social Impact of Dataset](#social-impact-of-dataset) - [Discussion of Biases](#discussion-of-biases) - [Other Known Limitations](#other-known-limitations) - [Additional Information](#additional-information) - [Dataset Curators](#dataset-curators) - [Licensing Information](#licensing-information) - [Citation Information](#citation-information) - [Contributions](#contributions) ## Dataset Description - **Homepage:** [QAngaroo](http://qangaroo.cs.ucl.ac.uk/) - **Repository:** [If the dataset is hosted on github or has a github homepage, add URL here]() - **Paper:** [Constructing Datasets for Multi-hop Reading Comprehension Across Documents](https://arxiv.org/abs/1710.06481) - **Leaderboard:** [leaderboard](http://qangaroo.cs.ucl.ac.uk/leaderboard.html) - **Point of Contact:** [Johannes Welbl](j.welbl@cs.ucl.ac.uk) ### Dataset Summary [More Information Needed] ### Supported Tasks and Leaderboards [More Information Needed] ### Languages [More Information Needed] ## Dataset Structure ### Data Instances [More Information Needed] ### Data Fields [More Information Needed] ### Data Splits [More Information Needed] ## Dataset Creation ### Curation Rationale [More Information Needed] ### Source Data [More Information Needed] #### Initial Data Collection and Normalization [More Information Needed] #### Who are the source language producers? [More Information Needed] ### Annotations [More Information Needed] #### Annotation process [More Information Needed] #### Who are the annotators? [More Information Needed] ### Personal and Sensitive Information [More Information Needed] ## Considerations for Using the Data ### Social Impact of Dataset [More Information Needed] ### Discussion of Biases [More Information Needed] ### Other Known Limitations [More Information Needed] ## Additional Information ### Dataset Curators [More Information Needed] ### Licensing Information [More Information Needed] ### Citation Information [More Information Needed] ### Contributions Thanks to [@patil-suraj](https://github.com/patil-suraj) for adding this dataset.
annotations_creators: - 众包(crowdsourced) language_creators: - 专家生成(expert-generated) language: - 英语 license: - CC BY-SA 3.0 multilinguality: - 单语种(monolingual) size_categories: - 1000 < 样本量 < 10000 source_datasets: - 原创 task_categories: - 问答(question-answering) task_ids: - 抽取式问答(extractive-qa) paperswithcode_id: medhop pretty_name: MedHop tags: - 多跳(multi-hop) dataset_info: - config_name: original features: - 字段名:id,数据类型:字符串 - 字段名:问题(query),数据类型:字符串 - 字段名:答案(answer),数据类型:字符串 - 字段名:候选答案集(candidates),类型:字符串序列 - 字段名:支持文档集(supports),类型:字符串序列 splits: - 拆分集名称:train(训练集),字节数:93937322,样本数:1620 - 拆分集名称:validation(验证集),字节数:16461640,样本数:342 download_size: 339843061 dataset_size: 110398962 - config_name: masked features: - 字段名:id,数据类型:字符串 - 字段名:问题(question),数据类型:字符串 - 字段名:答案(answer),数据类型:字符串 - 字段名:候选答案集(candidates),类型:字符串序列 - 字段名:支持文档集(supports),类型:字符串序列 splits: - 拆分集名称:train(训练集),字节数:95813584,样本数:1620 - 拆分集名称:validation(验证集),字节数:16800570,样本数:342 download_size: 339843061 dataset_size: 112614154 --- # MedHop 数据集卡片 ## 目录 - [数据集描述](#dataset-description) - [数据集概述](#dataset-summary) - [支持任务与排行榜](#supported-tasks-and-leaderboards) - [语言](#languages) - [数据集结构](#dataset-structure) - [数据实例](#data-instances) - [数据字段](#data-fields) - [数据拆分](#data-splits) - [数据集创建](#dataset-creation) - [数据集整理依据](#curation-rationale) - [源数据](#source-data) - [注释](#annotations) - [个人与敏感信息](#personal-and-sensitive-information) - [数据集使用注意事项](#considerations-for-using-the-data) - [数据集的社会影响](#social-impact-of-dataset) - [偏差讨论](#discussion-of-biases) - [其他已知局限性](#other-known-limitations) - [附加信息](#additional-information) - [数据集整理者](#dataset-curators) - [许可证信息](#licensing-information) - [引用信息](#citation-information) - [贡献声明](#contributions) ## 数据集描述 - **主页**:[QAngaroo](http://qangaroo.cs.ucl.ac.uk/) - **代码仓库**:[若数据集托管于GitHub或拥有GitHub主页,请在此处添加URL]() - **相关论文**:[跨文档多跳阅读理解数据集构建](https://arxiv.org/abs/1710.06481) - **排行榜**:[排行榜](http://qangaroo.cs.ucl.ac.uk/leaderboard.html) - **联系方式**:[约翰内斯·韦尔布(Johannes Welbl)](j.welbl@cs.ucl.ac.uk) ### 数据集概述 [More Information Needed] ### 支持任务与排行榜 [More Information Needed] ### 语言 [More Information Needed] ## 数据集结构 ### 数据实例 [More Information Needed] ### 数据字段 [More Information Needed] ### 数据拆分 [More Information Needed] ## 数据集创建 ### 数据集整理依据 [More Information Needed] ### 源数据 [More Information Needed] #### 初始数据收集与标准化 [More Information Needed] #### 源语言生成者是谁? [More Information Needed] ### 注释 [More Information Needed] #### 注释流程 [More Information Needed] #### 注释者是谁? [More Information Needed] ### 个人与敏感信息 [More Information Needed] ## 数据集使用注意事项 ### 数据集的社会影响 [More Information Needed] ### 偏差讨论 [More Information Needed] ### 其他已知局限性 [More Information Needed] ## 附加信息 ### 数据集整理者 [More Information Needed] ### 许可证信息 [More Information Needed] ### 引用信息 [More Information Needed] ### 贡献声明 感谢 [@patil-suraj](https://github.com/patil-suraj) 为本数据集的收录提供支持。
数据集卡片 MedHop
数据集描述
数据集摘要
- annotations_creators: 众包
- language_creators: 专家生成
- language: 英语
- license: CC BY-SA 3.0
- multilinguality: 单语
- size_categories: 1K<n<10K
- source_datasets: 原始数据
- task_categories: 问答
- task_ids: 抽取式问答
- paperswithcode_id: medhop
- pretty_name: MedHop
- tags: 多跳
数据集结构
配置
-
config_name: original
- features:
- id: 字符串
- query: 字符串
- answer: 字符串
- candidates: 字符串序列
- supports: 字符串序列
- splits:
- train:
- num_bytes: 93937322
- num_examples: 1620
- validation:
- num_bytes: 16461640
- num_examples: 342
- train:
- download_size: 339843061
- dataset_size: 110398962
- features:
-
config_name: masked
- features:
- id: 字符串
- question: 字符串
- answer: 字符串
- candidates: 字符串序列
- supports: 字符串序列
- splits:
- train:
- num_bytes: 95813584
- num_examples: 1620
- validation:
- num_bytes: 16800570
- num_examples: 342
- train:
- download_size: 339843061
- dataset_size: 112614154
- features:



