vikhram-labs/IndicConBench
收藏资源简介:
IndicConBench是一个针对印度的多语言宪法推理基准数据集,具有出版级和研究质量。它专门设计用于评估大语言模型在印度宪法上的事实准确性、推理能力、结构检索能力和法律智能,支持英语和印地语两种语言。该数据集由Vikhram S开发,Vikhram Labs发布,作为多语言法律AI领域的严格科学工具。它覆盖了465篇文章,包含8408个并行示例,总令牌数估计为1050651,并分为训练集(6708个示例)、验证集(832个示例)、测试集(848个示例)和GPQA集(20个示例,由专家编写的多项选择题)。数据集包含9个不同的任务家族,包括宪法问答、文章检索、宪法推理、摘要、蕴含、多跳推理、原则分类、文章链接和IndicConBench-GPQA,旨在全面评估法律AI的各个方面。数据集采用文章级分割以防止数据泄漏,并提供了科学的评估协议。
IndicConBench is a publication-grade, research-quality multilingual constitutional reasoning benchmark for India. It is specifically designed to evaluate the factual accuracy, reasoning capacity, structural retrieval capability, and legal intelligence of Large Language Models (LLMs) on the Constitution of India in both English and Hindi. Developed by Vikhram S and published under Vikhram Labs, this dataset serves as a rigorous scientific instrument for the multilingual legal AI domain. It covers 465 articles, with 8408 parallel examples, an estimated total of 1050651 tokens, and splits into train (6708 examples), validation (832 examples), test (848 examples), and GPQA (20 expert-authored multiple-choice examples). The dataset includes 9 diverse task families such as Constitutional QA, Article Retrieval, Constitutional Reasoning, Summarization, Entailment, Multi-hop Reasoning, Principle Classification, Article Linking, and IndicConBench-GPQA, aimed at evaluating distinct aspects of legal AI. It employs article-level partitioning to prevent data leakage and provides a scientific evaluation protocol.




