deu05232/promptriever-RQ1-same_version_seed12
收藏资源简介:
该数据集是一个用于信息检索或问答任务的结构化数据集,包含查询及其相关和不相关的文档段落。每个样本包括查询ID、查询文本、正相关段落列表(含文档ID、解释、评分、联合ID、文本和标题)、负相关段落列表(含文档ID、文本和标题)、仅指令字段、仅查询字段、是否含指令的布尔标志,以及新负例列表(结构与正相关段落类似)。数据集旨在支持模型训练,以区分查询与相关文档的匹配程度,适用于检索增强生成或文档排序等应用。
This dataset is a structured dataset for information retrieval or question-answering tasks, containing queries along with relevant and non-relevant document passages. Each sample includes a query ID, query text, a list of positive passages (with document ID, explanation, score, joint ID, text, and title), a list of negative passages (with document ID, text, and title), an only-instruction field, an only-query field, a boolean flag indicating whether instructions are present, and a list of new negatives (with a structure similar to positive passages). The dataset is designed to support model training for distinguishing the relevance of documents to queries, suitable for applications such as retrieval-augmented generation or document ranking.



