Hindi-BEIR
收藏资源简介:
Hindi-BEIR是一个针对印地语的大型检索基准数据集,由印度理工学院巴特那分校和IBM研究共同创建。该数据集包含15个多样化的子数据集,跨越8个不同的任务和5个领域,总计超过2700万份文档和近20万个查询。数据集的创建过程包括翻译现有英语数据集、从现有数据集中创建检索数据集以及编译多语言数据集。Hindi-BEIR旨在评估和推进印地语信息检索模型的性能,特别是在处理不同领域和任务的多样性方面。
Hindi-BEIR is a large-scale retrieval benchmark dataset for the Hindi language, jointly created by the Indian Institute of Technology Patna and IBM Research. This dataset includes 15 diverse sub-datasets spanning 8 distinct tasks and 5 domains, with a total of over 27 million documents and nearly 200,000 queries. The dataset construction process includes translating existing English datasets, creating retrieval datasets from existing datasets, and compiling multilingual datasets. Hindi-BEIR aims to evaluate and advance the performance of Hindi information retrieval models, particularly in handling diversity across different domains and tasks.

- 1Hindi-BEIR : A Large Scale Retrieval Benchmark in Hindi印度理工学院巴特那分校 · 2024年



