BeIR/hotpotqa-qrels
收藏资源简介:
BEIR Benchmark是一个异构基准,由18个不同的数据集组成,代表了9种信息检索任务,包括事实核查、问答、生物医学信息检索、新闻检索、论点检索、重复问题检索、引用预测、推文检索和实体检索。所有数据集均为英文,并已预处理,可用于实验。数据集结构包括corpus、queries和qrels文件,格式为jsonl和tsv。
BEIR Benchmark is a heterogeneous benchmark composed of 18 distinct datasets, representing 9 information retrieval tasks, including fact checking, question answering, biomedical information retrieval, news retrieval, argument retrieval, duplicate question retrieval, citation prediction, tweet retrieval, and entity retrieval. All datasets are in English and have been preprocessed for experimental use. The dataset structure includes corpus, queries, and qrels files, available in jsonl and tsv formats.
数据集概述
名称: BEIR Benchmark
语言: 英语 (en)
许可证: CC-BY-SA-4.0
多语言性: 单语
大小:
- MSMARCO: 1M<n<10M
- TREC-COVID: 100k<n<1M
- NFCorpus: 1K<n<10K
- NQ: 1M<n<10M
- HotpotQA: 1M<n<10M
- FiQA: 10K<n<100K
- ArguAna: 1K<n<10K
- Touche-2020: 100K<n<1M
- CQADupstack: 100K<n<1M
- Quora: 100K<n<1M
- DBpedia: 1M<n<10M
- SCIDOCS: 10K<n<100K
- FEVER: 1M<n<10M
- Climate-FEVER: 1M<n<10M
- SciFact: 1K<n<10K
任务类别:
- 文本检索
- 零样本检索
- 信息检索
- 零样本信息检索
任务ID:
- 段落检索
- 实体链接检索
- 事实检查检索
- 推文检索
- 引用预测检索
- 重复问题检索
- 论证检索
- 新闻检索
- 生物医学信息检索
- 问答检索
数据集结构
数据实例:
- 语料库:
.jsonl文件,包含文档的唯一标识符、标题和文本。 - 查询:
.jsonl文件,包含查询的唯一标识符和文本。 - qrels:
.tsv文件,包含查询ID、文档ID和分数。
数据集创建
许可证信息: CC-BY-SA-4.0
引用信息:
@inproceedings{ thakur2021beir, title={{BEIR}: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models}, author={Nandan Thakur and Nils Reimers and Andreas R{"u}ckl{e} and Abhishek Srivastava and Iryna Gurevych}, booktitle={Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2)}, year={2021}, url={https://openreview.net/forum?id=wCu6T5xFjeJ} }




