LETOR 4.0
收藏资源简介:
LETOR 4.0是由微软亚洲研究院创建的基准数据集,专注于学习排序研究。该数据集基于Gov2网页集合和TREC 2007及2008的Million Query轨道查询集,包含约1700个标记文档的MQ2007查询和约800个标记文档的MQ2008查询。数据集创建过程中采用了5折交叉验证策略,并提供了多种版本的处理数据。LETOR 4.0主要应用于搜索引擎的排序算法优化,旨在提高查询结果的相关性和准确性。
LETOR 4.0 is a benchmark dataset created by Microsoft Research Asia, focusing on learning-to-rank research. The dataset is based on the Gov2 web corpus and the Million Query track query sets from TREC 2007 and 2008. It includes approximately 1700 labeled documents for MQ2007 queries and around 800 labeled documents for MQ2008 queries. A 5-fold cross-validation strategy was adopted during the dataset's construction, and multiple processed versions of the data are provided. LETOR 4.0 is mainly applied to the optimization of ranking algorithms for search engines, with the goal of enhancing the relevance and accuracy of query results.




