遇见数据集

yena0rganization/global-llm-value-gap

收藏
Hugging Face2026-04-28 更新2026-05-03 收录
官方服务:

资源简介:

该数据集是一个基于World Values Survey(WVS) Wave 7构建的模拟响应数据集,旨在测量大型语言模型(LLM)与人类群体之间的价值观差异。数据集通过66个国家的人口统计分布反映分层抽样角色(persona),并在6个LLM模型(GPT-4o, GPT-4o-mini, Gemini-2.5-flash, Gemini-2.0-flash, DeepSeek-reasoner, DeepSeek-chat)上收集对23个伦理问题的回答,与真实WVS人类回答进行系统比较。数据集包含910,800个LLM响应,覆盖9个伦理类别(基于道德基础理论),如不诚实与腐败、性伦理、生命伦理等。每个国家的角色(persona)数量为100个,通过分层抽样方式生成,并经过TOST检验验证其统计等效性。

This dataset is a simulated response dataset based on World Values Survey (WVS) Wave 7, designed to measure the value gap between large language models (LLMs) and human groups. The dataset reflects the demographic distribution of 66 countries through stratified sampling personas and collects responses to 23 ethical questions from 6 LLM models (GPT-4o, GPT-4o-mini, Gemini-2.5-flash, Gemini-2.0-flash, DeepSeek-reasoner, DeepSeek-chat), systematically comparing them with real WVS human responses. The dataset includes 910,800 LLM responses, covering 9 ethical categories (based on moral foundation theory), such as dishonesty & corruption, sexual ethics, bioethics, etc. Each country has 100 personas generated through stratified sampling, validated for statistical equivalence via TOST tests.

提供机构:
yena0rganization
二维码
社区交流群
二维码
科研交流群
商业服务