遇见数据集

Sanskriti

收藏
魔搭社区2026-08-23 更新2026-08-23 收录
官方服务:

资源简介:

*** # SANSKRITI: A Comprehensive Benchmark for Evaluating Language Models' Knowledge of Indian Culture **Link:** [arxiv.org/abs/2506.15355](arxiv.org/abs/2506.15355) ## Dataset Description The **SANSKRITI** benchmark is the largest dataset created to evaluate Language Models' (LMs) comprehension and reasoning capabilities regarding the rich cultural diversity of India.It addresses the critical need for culturally-aware benchmarks, as the global effectiveness of LMs depends on their understanding of local socio-cultural contexts. ### Dataset Composition SANSKRITI comprises **21,853 meticulously curated question-answer pairs**. ***Geographical Coverage:** It spans all **28 states and 8 union territories** of India. ***Cultural Attributes:** The dataset covers **sixteen key attributes** of Indian culture, providing a comprehensive representation of India's cultural tapestry: * Rituals and ceremonies * History * Tourism * Cuisine * Dance and music * Costume * Language * Art * Festivals * Religion * Medicine * Transport * Sports * Nightlife * Personalities *** ## Data Distribution The dataset's questions are systematically categorized into four types and show a balanced distribution across these categories: ### Question Types | Question Type | Description | | :--- | :--- | | **Country Prediction** | Determine the country (e.g., India) based on a cultural statement. | | **Association Based** | Identify the cultural context or entity most closely associated with a cultural element. | | **General Awareness** | Factual questions about general cultural attributes. | | **State Prediction** | Identify the specific Indian state referenced in a cultural statement. | ![{C29E2927-A4B9-4429-9775-27D9B3E04373}](https://cdn-uploads.huggingface.co/production/uploads/66b725125ce4c02bda63e7f0/4HVQH02mdmdZ-bh5Weo_w.png) ### Geographical Distribution ![{08D7BC8E-315B-413E-ABD1-E6B807B72409}](https://cdn-uploads.huggingface.co/production/uploads/66b725125ce4c02bda63e7f0/PNm20O2OIQCkLcQ5shsip.png) *** ## Evaluation and Findings The SANSKRITI benchmark was used to evaluate leading Large Language Models (**LLMs**), Small Language Models (**SLMs**), and Indic Language Models (**ILMs**). The evaluation was performed using a **zero-shot, multiple-choice question (MCQ)** format, with accuracy as the sole metric. ### Key Performance Trends * **Overall Top Performer:** **GPT-4o** demonstrated the best overall performance. * **Open-Source LLMs:** **LLAMA-3.1-70B-Instruct** achieved the highest accuracy (0.86) among open-source LLMs. * **SLMs:** **Qwen2-1.5B-Instruct** emerged as the best-performing SLM, with a score of 0.74, highlighting that some SLMs can possess better domain-specific knowledge than even larger models. ![{18C2FBEB-D5B9-481A-9EF6-CA8199207222}](https://cdn-uploads.huggingface.co/production/uploads/66b725125ce4c02bda63e7f0/HORDo3903vZByc2JJjT4D.png) * **Question Type Performance:** * Models achieved the **highest accuracy on General Awareness** questions. * Models showed the **lowest accuracy on State Prediction** questions. * For **Association-Based** questions, **LLAMA-3.1-70B** outperformed all other models, including GPT-4o. * **Cultural Attribute Performance:** Models performed notably well on **Religion, Medicine, and Cultural Common Sense**, but generally **struggled with Costumes, Cuisines, and Art**. * **Geographical Struggles:** LMs significantly struggled with questions pertaining to **North-Eastern states (e.g., Sikkim, Arunachal Pradesh, Tripura)**, as well as **Bihar and Jharkhand**. They performed better for states with globally recognized cities like **Delhi and Maharashtra**. *** ## Dataset Creation and Methodology ### Data Sourcing and Organization The data was sourced from multiple diverse and reliable platforms to ensure authenticity and depth, some of the most used sources are listed below: 1. Wikipedia 2. Ritiriwaz (for customs and rituals) 3. Holidify (for travel, festivals, cuisines, and landmarks) 4. Arts and Culture 5. Times of India The information was organized in a structured format: ("state name": "attribute": "scrapped data related to the attribute and state"). ### Annotation Process A team of **40 annotators** — including native and bilingual speakers from various Indian states — was employed to create and validate the questions. * The annotators were divided into four specialized sub-teams, each focusing on one of the four question types. * A rigorous **cross-validation** process was used, where outputs from each sub-team were reviewed by another sub-team to resolve ambiguities and ensure accuracy. * **Cultural sensitivity and ethical guidelines** were prioritized, emphasizing the avoidance of stereotypes and encouraging respectful, inclusive framing. * Annotators were compensated for both question creation and verification. *** ## Limitations and Future Work The authors acknowledge several limitations of the current SANSKRITI dataset: 1.**Limited Scope of Cultural Attributes:** Only sixteen attributes are covered, not fully representing the entirety of Indian culture. Future work plans to expand attributes to include things like regional folklore and ecological heritage. 2.**Lack of State-Specific Multilingual Queries:** The current dataset lacks state-specific multilingual questions. This is a key priority for immediate extension. 3.**Absence of Visual Question-Answering (VQA) Tasks:** The dataset does not include VQA tasks. Future work plans to build a visual question-answering dataset and a multilingual version. 4.**Lack of Diversity in Questions:** Questions are multiple-choice and fact-based, not explicitly requiring complex reasoning or causal understanding. 📂 Citation If you use this dataset, please cite: ```bibtex @inproceedings{maji2025sanskriti, title={SANSKRITI: A Comprehensive Benchmark for Evaluating Language Models’ Knowledge of Indian Culture}, author={Maji, Arijit and Kumar, Raghvendra and Ghosh, Akash and Anushka and Saha, Sriparna}, booktitle={Findings of the Association for Computational Linguistics: ACL 2025}, year={2025} } ```

提供机构:
maas
创建时间:
2026-08-20
二维码
社区交流群
二维码
科研交流群
商业服务