BhashaBench V1
收藏资源简介:
BhashaBench V1是一个全面的、针对印度特定知识系统的多任务、双语文本数据集,用于评估大型语言模型(LLMs)。该数据集包含74166个经过精心策划的问题和答案对,其中52494个用英语编写,21672个用印地语编写,来源于真实的政府和特定领域的考试。数据集涵盖了农业、法律、金融和印度传统医学四大主要领域,包括90多个子领域和500多个主题,支持对模型在印度多样知识领域的深入评估。该数据集的创建旨在解决印度特定领域和语言中现有基准的局限性,为评估模型在不同领域的文化理解和专业知识提供了一种工具。
BhashaBench V1 is a comprehensive, multi-task bilingual text dataset targeting India-specific knowledge systems for evaluating Large Language Models (LLMs). It contains 74,166 carefully curated question-answer pairs, with 52,494 written in English and 21,672 in Hindi, sourced from real-world government and domain-specific examinations. The dataset covers four core domains: agriculture, law, finance, and Ayurveda, encompassing over 90 sub-domains and more than 500 topics, which enables in-depth assessment of models' capabilities across diverse knowledge fields specific to India. This dataset was developed to address the limitations of existing benchmarks in India-specific domains and languages, providing a specialized tool for evaluating models' cultural understanding and professional expertise across various fields.
BHASHABENCH V1 数据集概述
数据集基本信息
- 数据集名称: BHASHABENCH V1
- 数据规模: 74,166个精心策划的问答对
- 语言分布:
- 英语: 52,494个问题(70.8%)
- 印地语: 21,672个问题(29.2%)
核心领域覆盖
主要领域
- 农业(Agriculture)
- 法律(Legal)
- 金融(Finance)
- 阿育吠陀(Ayurveda)
细分范围
- 90+个子领域
- 500+个主题
- 支持细粒度评估
任务格式
- 多项选择题
- 断言推理题
- 填空题
- 阅读理解题
- 超过90%为多项选择题
数据来源与质量
数据收集
- 来自国家和州级评估的真实考试材料
- 涵盖40多种不同类型的考试
- 包括:国家竞争性考试、领域特定学位考试、专业认证考试、州级公务员考试
数据处理
- 使用Surya OCR进行高质量文本提取
- 基于GPT的系统将内容结构化为标准JSON问答对
- 经过广泛的清理和去重
- 严格的专家验证确保准确性和文化相关性
难度级别
- 简单(Easy)
- 中等(Medium)
- 困难(Hard)
评估结果概况
- 评估了29+个大型语言模型
- 模型在不同领域和语言间存在显著性能差距
- 模型在所有领域中英语内容的表现均优于印地语
- 具体表现示例:GPT-4o在法律领域准确率达76.49%,在阿育吠陀领域仅为59.74%
资源获取
- 所有代码、基准和资源均公开可用
- 支持开放研究

- 1BhashaBench V1: A Comprehensive Benchmark for the Quadrant of Indic Domains印度理工学院孟买分校(印度理工学院孟买分校) · 2025年



