mMARCO
收藏资源简介:
mMARCO是一个多语言版本的MS MARCO段落排序数据集,由坎皮纳斯大学神经网络实验室创建,包含13种语言。该数据集通过机器翻译创建,旨在解决非英语语言在信息检索任务中训练资源稀缺的问题。数据集包含超过53万条查询-段落相关对,适用于训练和评估深度学习模型。mMARCO的创建不仅丰富了多语言信息检索的训练资源,还通过零样本学习场景的评估,展示了其对提升模型效果的潜力。
mMARCO is a multilingual variant of the MS MARCO passage ranking dataset, developed by the Neural Networks Lab at the University of Campinas and covering 13 languages. Constructed via machine translation, this dataset aims to address the scarcity of training resources for non-English languages in information retrieval tasks. It contains over 530,000 query-passage relevance pairs, suitable for training and evaluating deep learning models. Beyond enriching the training resources for multilingual information retrieval, mMARCO also demonstrates its potential to enhance model performance through evaluations in zero-shot learning scenarios.

- 1mMARCO: A Multilingual Version of the MS MARCO Passage Ranking Dataset坎皮纳斯大学神经网络实验室 · 2022年



