遇见数据集

anonymous-submission-nips/mlaire-belebele

收藏
Hugging Face2026-05-06 更新2026-05-31 收录
官方服务:

资源简介:

MLAIRE-BELEBELE是一个用于语言感知检索评估的数据集,基于Belebele数据集重新格式化。它包含488个底层段落,每个段落提供122种语言版本(通过原始`link`字段全局连接)。相关性通过`group_id`匹配进行编码。该数据集是MLAIRE基准的一部分,已匿名提交至NeurIPS 2026 Evaluations & Datasets Track。数据布局包括corpus、queries和qrels,分别存储文档、查询和相关关系,支持语言感知指标(如LPR、Lang-nDCG、Lang-Recall)和4-way top-1失败分解(perfect / lang_fail / sem_fail / both_fail)的计算。

MLAIRE-BELEBELE is a dataset for language-aware retrieval evaluation, reformatted based on the Belebele dataset. It contains 488 underlying paragraphs, each with 122 language versions that are globally linked via the original `link` field. Relevance is encoded through `group_id` matching. This dataset is part of the MLAIRE benchmark, and has been anonymously submitted to the NeurIPS 2026 Evaluations & Datasets Track. Its data layout includes corpus, queries, and qrels, which store documents, queries, and relevance relationships respectively, supporting the calculation of language-aware metrics such as LPR, Lang-nDCG, Lang-Recall, and 4-way top-1 failure decomposition (perfect / lang_fail / sem_fail / both_fail).

二维码
社区交流群
二维码
科研交流群
商业服务