Kukedlc/Personas-Argentinas-1k
收藏资源简介:
Personas Argentinas是一个多用途的合成数据集,包含1000个阿根廷人的合成人口,基于官方统计来源(如阿根廷国家统计和普查局INDEC的永久家庭调查EPH、2022年全国人口普查、Latinobarómetro 2023阿根廷民意调查等)构建,具有统计严谨性。每个合成人物整合了以下信息:基于人口普查和官方调查锚定的人口统计学和社会经济特征(如年龄、性别、教育水平、职业、收入、地区等);基于大五人格模型(OCEAN)的个性档案;通过分层统计插值方法基于真实民意调查数据校准的意见、价值观和意识形态立场(如左-右意识形态、投票意向、对机构的信任度等);以及连贯的传记(包括生活史、区域语言风格、兴趣、目标等)。数据集旨在支持可重复研究、社会模拟、人口建模、合成调查和焦点小组、受众细分测试,以及具有个性和对话的智能体开发。数据以JSON Lines和Parquet格式提供,包含71列。数据集是持续扩展的第一个版本,主要覆盖城市人口,并注意了内部一致性和多样性控制。
Personas Argentinas is a multipurpose synthetic dataset comprising a synthetic population of 1000 Argentine individuals, constructed with statistical rigor from official sources. Each synthetic person integrates demographics anchored in the national census and official surveys, a personality profile, an opinion and values positioning calibrated against real data, and a coherent biography. It is designed for reproducible research, social simulation, and population modeling. The dataset is based on official statistical sources such as INDECs Permanent Household Survey (EPH), the 2022 National Census, Latinobarómetro 2023 Argentina public opinion survey, and others. It includes 71 columns covering demographics, personality (Big Five OCEAN traits), opinion and values (e.g., ideology, voting intention, institutional trust), and a dense biography. The data is provided in JSON Lines and Parquet formats. This is the first version, continuously expanding, with a focus on urban populations and controlled for internal coherence and diversity.




