遇见数据集

TREC Fair Ranking 2022 Single-Region Subset for Quantification and Fairness Estimation

收藏
Zenodo2026-08-14 更新2026-08-20 收录
官方服务:

资源简介:

Dataset TREC Fair # data points 1,152,808 data modality Text sensitive attribute subcontinental region |S| 21 |Y| 45 target variable topic relevance The TREC Fair dataset is derived from the TREC Fair Ranking 2022 collection and consists of English Wikipedia articles annotated with subcontinental geographic metadata. To support the study of representation across regions, only the single-label portion of the dataset is retained, that is, articles assigned to exactly one subcontinental region. The sensitive attribute corresponds to subcontinental region, with labels including categories such as Northern America, Southern Europe, and Eastern Asia. The target variable is topic relevance, operationalised through query-specific document sets used in the ranking task. Each document is represented by concatenating the article title and plain-text article content. For the sparse representation used in the main experiments, the text is transformed using TF-IDF vectorization with English stopword removal, unigram and bigram features, and a vocabulary capped at the 10,000 most frequent terms. The resulting collection contains 1,152,808 documents after preprocessing. The dataset is used in a ranking setting, where each query defines a topic-specific subset of relevant documents and fairness is evaluated with respect to the regional distribution induced by the ranking. Original dataset/source: @inproceedings{trec-fair-ranking-2021, Author = {Michael D. Ekstrand and Graham McDonald and Amifa Raj and Isaac Johnson}, Booktitle = {The Thirtieth Text REtrieval Conference (TREC 2021) Proceedings}, Title = {Overview of the TREC 2021 Fair Ranking Track}, Year = {2022} } This work has been funded by the QuaDaSh project “Finanziato dall’Unione europea- Next Generation EU, Missione 4 Componente 2 CUP B53D23026250001”

提供机构:
Zenodo
创建时间:
2026-08-14
二维码
社区交流群
二维码
科研交流群
商业服务