IndicRAGSuite
收藏资源简介:
IndicRAGSuite是一个针对印度语言RAG系统的数据集,包括IndicMSMarco,一个用于评估13种印度语言检索质量和响应生成的多语言基准。该数据集还包括从19种印度语言的维基百科中提取的约1400万个(问题、答案、相关段落)三元组。数据通过利用先进的语言模型生成,并包括MS MARCO数据集的翻译版本,以确保与现实世界的信息检索任务保持一致。
IndicRAGSuite is a dataset dedicated to Indian language Retrieval-Augmented Generation (RAG) systems. It includes IndicMSMarco, a multilingual benchmark for evaluating retrieval quality and response generation across 13 Indian languages. Additionally, the dataset contains approximately 14 million (question, answer, relevant passage) triples extracted from Wikipedia articles spanning 19 Indian languages. The data was generated using state-of-the-art language models, and incorporates translated versions of the MS MARCO dataset to ensure consistency with real-world information retrieval tasks.




