遇见数据集

envyr/scidocs_with_instructions

收藏
Hugging Face2026-05-25 更新2026-05-31 收录
官方服务:

资源简介:

该数据集是一个用于信息检索任务的数据集,包含三个主要部分:语料库(corpus)、查询(queries)和相关度评分(qrels)。语料库包含25,657个文档,每个文档有ID、标题和文本内容;查询部分包含1,000个查询,每个查询有ID和文本;相关度评分部分包含29,928个评分项,用于关联查询与文档,并提供相关性分数。数据集适用于训练和评估检索模型,如文档排序或问答系统。

This dataset is designed for information retrieval tasks and consists of three main components: a corpus, queries, and relevance judgments (qrels). The corpus includes 25,657 documents, each with an ID, title, and text content. The queries section contains 1,000 queries, each with an ID and text. The qrels section includes 29,928 relevance judgments that link queries to documents with a score. It is suitable for training and evaluating retrieval models, such as document ranking or question-answering systems.

提供机构:
envyr
二维码
社区交流群
二维码
科研交流群
商业服务