遇见数据集

deu05232/repro_msmarco-w-instructions_seed42-multipos-subset_replace_version

收藏
Hugging Face2026-05-25 更新2026-05-31 收录
官方服务:

资源简介:

该数据集是一个用于信息检索或问答系统任务的结构化数据集,包含查询和相关的段落样本。主要特征包括:查询ID和查询文本;正相关段落列表,每个段落包含文档ID、解释文本、FollowIR评分、联合ID、文本内容和标题;负相关段落列表,每个段落包含文档ID、文本内容和标题;仅指令文本字段;仅查询文本字段;布尔值指示是否包含指令;以及新负例段落列表,结构与正相关段落类似。数据集分为训练集,包含978,432个样本,总大小约为9.66 GB。数据可能用于训练检索模型或评估查询与文档的相关性。

This dataset is a structured dataset for information retrieval or question-answering tasks, containing queries and associated passage samples. Key features include: query ID and query text; a list of positive passages, each with document ID, explanation text, FollowIR score, joint ID, text content, and title; a list of negative passages, each with document ID, text content, and title; an only instruction text field; an only query text field; a boolean indicating whether instructions are included; and a list of new negative passages, with a structure similar to positive passages. The dataset is split into a training set with 978,432 examples and a total size of approximately 9.66 GB. It may be used for training retrieval models or evaluating query-document relevance.

提供机构:
deu05232
二维码
社区交流群
二维码
科研交流群
商业服务