IndiCASA
收藏资源简介:
IndiCASA数据集是一份针对印度社会文化背景下的大型语言模型(LLM)偏见评估的新颖数据集。该数据集包含2,575个人工验证的句子,涵盖了五个社会人口学维度:种姓、性别、宗教、残疾和社会经济地位。IndiCASA数据集旨在帮助LLM更好地理解印度社会文化中的细微偏见,从而提高模型的公平性和准确性。数据集的创建过程采用了人类专家与人工智能协作的方法,确保了数据的质量和多样性。IndiCASA数据集可用于评估LLM在印度语境下的偏见程度,并帮助开发者改进模型的公平性和文化敏感性。
The IndiCASA dataset is a novel dataset for bias evaluation of Large Language Models (LLMs) in the context of Indian socio-cultural settings. It contains 2,575 human-validated sentences covering five sociodemographic dimensions: caste, gender, religion, disability, and socioeconomic status. The IndiCASA dataset aims to help LLMs better understand subtle biases in Indian socio-cultural contexts, thereby improving model fairness and accuracy. The dataset was created using a human-AI collaborative methodology to ensure data quality and diversity. The IndiCASA dataset can be used to evaluate the degree of bias in LLMs within the Indian context, and assist developers in enhancing model fairness and cultural sensitivity.




