INDIC QA BENCHMARK
收藏资源简介:
INDIC QA BENCHMARK是由印度孟买理工学院和IBM印度研究院共同创建的多语言问题回答评估基准,涵盖11种主要印度语言。该数据集包括提取式和生成式问题回答任务,涉及多个领域如地理、印度文化、新闻等。数据集的创建过程包括翻译现有数据集和使用Gemini模型生成合成数据,并通过人工验证确保质量。该数据集主要用于评估和提升大型语言模型在低资源印度语言中的问题回答能力。
INDIC QA BENCHMARK is a multilingual question answering evaluation benchmark co-created by the Indian Institute of Technology Bombay and IBM Research India, covering 11 major Indian languages. This dataset includes extractive and generative question answering tasks, spanning multiple domains such as geography, Indian culture, news and others. The dataset was developed by translating existing datasets and generating synthetic data using the Gemini model, with manual validation conducted to ensure data quality. It is primarily used to evaluate and enhance the question answering capabilities of large language models (LLMs) in low-resource Indian languages.
chaii - 印地语和泰米尔语问答数据集
概述
该数据集旨在识别印度语言文章中的问题答案。
详细描述
- 标题: chaii - Hindi and Tamil Question Answering
- 描述: 识别印度语言文章中的问题答案

- 1INDIC QA BENCHMARK: A Multilingual Benchmark to Evaluate Question Answering capability of LLMs for Indic Languages印度孟买理工学院, IBM印度研究院 · 2024年



