遇见数据集

Rank-DistiLLM Novelty

收藏
Zenodo2025-04-02 更新2026-05-26 收录
官方服务:

资源简介:

This dataset contains the novelty ranking data for the paper Set-Encoder: Permutation-Invariant Inter-Passage Attention for Listwise Passage Re-Ranking with Cross-Encoders. The dataset is based off the Rank-Distillm dataset (Paper, Dataset). The two run files contain the top 100 retrieved passages for 10k queries from the MS MARCO passage training dataset re-ranked by the RankZephyr model for BM25 and ColBERTv2, respectively. The central difference to the Rank-DistiLLM dataset files is that the run files in this dataset are additionally grouped by lexical similarity. All passages within a ranking a clustered according to their jaccard similarity. Any passages with a jaccard similarity > 0.5 are grouped into a cluster and have the same id in the second column (the Q0 column in the standard TREC run format).The two qrel files are equivalent to the official TREC Deep Learning 2019 and 2020 qrels, but are also clustered by their lexical similarity (following the strategy outlined above) to enable evaluation of novelty-aware models.

提供机构:
Zenodo
创建时间:
2025-04-02
二维码
社区交流群
二维码
科研交流群
商业服务