BEIR-NL
收藏资源简介:
BEIR-NL是一个用于荷兰语信息检索的零样本评估基准,由安特卫普大学CLiPS团队通过自动翻译BEIR数据集中的14个子数据集创建。该数据集涵盖了从生物医学到金融等多个领域的信息检索任务,包含大量查询和文档,平均每个查询对应多个相关文档。数据集的创建过程包括选择合适的翻译工具(如Gemini-1.5flash)进行批量翻译,并进行了翻译质量评估。BEIR-NL旨在为荷兰语信息检索模型的开发和评估提供基础,解决荷兰语在信息检索研究中资源匮乏的问题。
BEIR-NL is a zero-shot evaluation benchmark for Dutch information retrieval, developed by the CLiPS team at the University of Antwerp through automatic translation of 14 sub-datasets from the original BEIR dataset. This dataset encompasses information retrieval tasks across diverse domains spanning from biomedicine to finance, and includes a substantial corpus of queries and documents, with an average of multiple relevant documents per query. The process of creating BEIR-NL involved selecting appropriate translation tools (such as Gemini-1.5 Flash) for batch translation, followed by translation quality evaluation. BEIR-NL is designed to provide a foundational resource for the development and evaluation of Dutch-language information retrieval models, addressing the issue of scarce Dutch-language resources in information retrieval research.

- 1BEIR-NL: Zero-shot Information Retrieval Benchmark for the Dutch Language安特卫普大学CLiPS · 2024年



