tomaarsen/zelo-preferences-10kx100-quantile-anchor
收藏资源简介:
该数据集是一个用于信息检索或排名任务的数据集,包含查询和文档对的相关性评分。数据集中每个示例包括查询ID、查询领域、查询文本、两个文档的索引、文档内容、总体评分、多个模型(如gpt-oss-20b、gemma-4-31b-it、granite-4.1-30b)的评分细节、有效标记数量以及教师标签。数据集旨在支持模型训练,特别是用于比较和评估不同模型在文档排名上的表现。它包含一个训练分割,共有1001000个示例,适用于大规模机器学习应用。
This dataset is designed for information retrieval or ranking tasks, containing relevance scores for query-document pairs. Each example includes a query ID, domain, query text, indices for two documents, document content, overall score, detailed scores from multiple models (e.g., gpt-oss-20b, gemma-4-31b-it, granite-4.1-30b), a validity count, and teacher labels. The dataset is intended for model training, particularly for comparing and evaluating the performance of different models on document ranking. It includes a training split with 1,001,000 examples, suitable for large-scale machine learning applications.



