Bharat_NanoArguAna_gu
收藏资源简介:
该数据集是Bharat-NanoBEIR集合的一部分,提供了印度语言的信息检索数据集。它是NanoArguAna数据集的Gujarati版本,专门用于信息检索任务。数据集包含三个主要部分:Corpus(文档集合)、Queries(搜索查询)和QRels(查询与相关文档的连接)。适用于Gujarati语言的信息检索系统开发、多语言搜索能力评估、跨语言信息检索研究以及Gujarati语言模型的搜索任务基准测试。
This dataset is a component of the Bharat-NanoBEIR collection, serving as an Indian-language information retrieval dataset. It is the Gujarati variant of the NanoArguAna dataset, specifically designed for information retrieval tasks. The dataset includes three core components: Corpus (document collection), Queries (search queries), and QRels (mappings between queries and their relevant documents). It is applicable to the development of Gujarati-language information retrieval systems, evaluation of multilingual search capabilities, cross-language information retrieval research, and benchmarking of search tasks for Gujarati language models.




