BEIR-PL
收藏资源简介:
BEIR-PL是一个专为波兰语设计的大型异构信息检索基准数据集,由弗罗茨瓦夫理工大学创建。该数据集包含13个子数据集,旨在促进现代波兰语模型的开发、训练和评估。数据集内容涵盖多种信息检索任务,如问题回答和实体链接,数据来源于多个开放资源。创建过程中,研究团队使用机器翻译技术将原始数据集翻译成波兰语,并进行了细致的评估和比较。BEIR-PL数据集的应用领域广泛,特别适用于零样本学习方法,为波兰语自然语言处理领域提供了重要的资源和基准。
BEIR-PL is a large-scale heterogeneous information retrieval benchmark dataset tailored specifically for Polish, created by Wrocław University of Science and Technology. It comprises 13 sub-datasets, with the goal of facilitating the development, training and evaluation of modern Polish language models. The dataset covers a diverse range of information retrieval tasks including question answering and entity linking, with data sourced from multiple open resources. During its creation, the research team utilized machine translation technologies to translate the original datasets into Polish, and conducted meticulous evaluation and comparative analysis. BEIR-PL has a wide range of application scenarios, and is particularly well-suited for zero-shot learning methods, providing essential resources and benchmarks for the field of Polish natural language processing.




