Pancasila-Dilemmas
收藏资源简介:
Pancasila-Dilemmas是由天津大学等机构构建的印尼价值观评估数据集,旨在评估大语言模型对印尼本土价值观的契合度。该数据集包含1,834个源自印尼新闻的多选题,涵盖宗教、人道、团结、民主和社会正义五大Pancasila核心价值观,每个问题由5名印尼本土标注者进行主观标注以捕捉人类偏好多样性。数据集通过爬取Kompas新闻门户、利用GPT-4o生成困境场景与选项,并经过母语者校对和质量控制流程精心构建。其核心应用于衡量大语言模型在印尼特定文化语境中的价值对齐能力,解决现有评估体系过度依赖西方或普世价值观而忽略地区性价值差异的问题。
Pancasila-Dilemmas is an Indonesian values assessment dataset developed by Tianjin University and other institutions, designed to evaluate the alignment between large language models (LLMs) and Indonesian indigenous values. This dataset contains 1,834 multiple-choice questions sourced from Indonesian news, covering the five core Pancasila values: religion, humanity, unity, democracy, and social justice. Each question was subjectively annotated by five local Indonesian annotators to capture the diversity of human preferences. The dataset was meticulously constructed through three main stages: scraping news content from the Kompas news portal, generating dilemma scenarios and corresponding options using GPT-4o, followed by native speaker proofreading and a strict quality control workflow. Its core application is to measure the value alignment capability of LLMs in the specific Indonesian cultural context, addressing the limitation that existing evaluation systems overly rely on Western or universal values while ignoring regional value differences.
数据集概述
数据集名称:Pancasila-Dilemmas
来源:该数据集来源于论文《Pancasila-Dilemmas: Evaluating Large Language Models on Indonesian Human Value Dilemmas Grounded in Pancasila》。
所属机构:tjunlp-lab(TJU NLP Lab)
数据内容:该数据集专注于评估大型语言模型在基于“潘查希拉”(Pancasila,印度尼西亚国家哲学原则)的印尼人类价值观困境中的表现。数据以价值观困境的形式呈现,用于测试和衡量语言模型对印尼本土文化价值观的理解与对齐程度。
主要用途:用于评估和测试大型语言模型在面对与潘查希拉原则相关的印尼人类价值观困境时的推理与判断能力,属于价值观对齐与跨文化评估领域的研究资源。
论文背景:该数据集是论文的配套资源,旨在为大语言模型在印尼文化语境下的伦理与价值观评估提供标准化的测试基准。




