百度-ULTR数据集
收藏资源简介:
百度-ULTR数据集是由百度公司和密歇根州立大学合作创建的大规模无偏学习排序数据集,包含12亿随机抽样的搜索会话和7008个专家标注查询(397,572个查询文档对)。该数据集提供了原始语义特征和不同大小的预训练语言模型,以及丰富的展示信息和用户反馈,如停留时间,支持多任务学习和用户参与度优化。数据集旨在解决现有数据集在语义特征提取、展示信息完整性和真实用户反馈方面的不足,推动无偏学习排序的研究。
The Baidu-ULTR Dataset is a large-scale unbiased learning-to-rank dataset jointly created by Baidu Inc. and Michigan State University. It includes 1.2 billion randomly sampled search sessions and 7008 expert-annotated queries, amounting to 397,572 query-document pairs in total. This dataset provides raw semantic features, pre-trained language models of varying sizes, as well as comprehensive display information and user feedback including dwell time, enabling multi-task learning and user engagement optimization. The dataset is intended to resolve the deficiencies of existing datasets in semantic feature extraction, integrity of display information and real-world user feedback, thus advancing research in the field of unbiased learning-to-rank.




