anonymous-submission-nips/mlaire-belebele
收藏资源简介:
MLAIRE-BELEBELE是一个用于语言感知检索评估的数据集,基于Belebele数据集重新格式化。它包含488个底层段落,每个段落提供122种语言版本(通过原始`link`字段全局连接)。相关性通过`group_id`匹配进行编码。该数据集是MLAIRE基准的一部分,已匿名提交至NeurIPS 2026 Evaluations & Datasets Track。数据布局包括corpus、queries和qrels,分别存储文档、查询和相关关系,支持语言感知指标(如LPR、Lang-nDCG、Lang-Recall)和4-way top-1失败分解(perfect / lang_fail / sem_fail / both_fail)的计算。
MLAIRE-BELEBELE is a dataset for language-aware retrieval evaluation, reformatted based on the Belebele dataset. It contains 488 underlying paragraphs, each with 122 language versions that are globally linked via the original `link` field. Relevance is encoded through `group_id` matching. This dataset is part of the MLAIRE benchmark, and has been anonymously submitted to the NeurIPS 2026 Evaluations & Datasets Track. Its data layout includes corpus, queries, and qrels, which store documents, queries, and relevance relationships respectively, supporting the calculation of language-aware metrics such as LPR, Lang-nDCG, Lang-Recall, and 4-way top-1 failure decomposition (perfect / lang_fail / sem_fail / both_fail).



