stratasynth-personas-8-countries
收藏资源简介:
StrataSynth Personas (8 Countries) 是一个包含10,000个合成人物的多语言数据集,覆盖西班牙、墨西哥、美国、英国、德国、法国、意大利和巴西八个国家,每个国家1,250个样本。数据集中每个人物都包含27个结构化字段,包括唯一ID、姓名、国家、语言、年龄(18-85岁)、性别、地区、城市规模、教育程度、职业(当前和上一份)、收入水平(低/中/高,基于国家实际分布)、家庭类型、婚姻状况、子女数量及年龄、住房类型、就业状况、是否与父母同住、日常作息、人物传记、说话风格、爱好与兴趣(5项)、技能(3-6项)、未来目标(1-3项)、消费者特征(购物风格、价格敏感度、品牌忠诚度、可持续关注度、技术采纳度)以及人格特征(大五人格:开放性、尽责性、外向性、宜人性、神经质,以及成人依恋的焦虑和回避维度,均为0-1连续值)。所有文本字段(日常作息、传记、说话风格、爱好、技能、目标)均使用人物所属国家的语言(英语、西班牙语、德语、法语、意大利语、葡萄牙语)。数据集的构建以真实人口统计为基础:年龄、地区、教育、收入、居住安排、家庭结构(以及五个国家的职业)均遵循各国成人人口的真实分布;姓名来自各国真实姓名频率;职业使用当地惯用名称;收入水平反映各国按教育和年龄划分的家庭收入分布。版本2中引入了真实职业(西班牙、法国、美国、墨西哥、巴西)和真实就业率。数据集适用于多种场景:在部署前使用数千个逼真用户测试聊天机器人和代理(客户服务、银行、健康、公共服务);进行公平性和偏差检查(向不同年龄、国家或性别的用户提出相同问题并比较回答);播种合成数据(生成多样化的指令);针对细分群体进行合成受访者预测试调查、产品和信息;以及角色扮演和角色工作(训练人物条件模型或游戏角色)。数据集以JSON Lines格式发布,每个国家一个独立文件(personas_es.jsonl等),可通过Hugging Face Datasets库按国家加载或合并为全部10,000个样本。许可协议为CC BY 4.0,使用时需注明出处为StrataSynth。
StrataSynth Personas (8 Countries) is a multilingual dataset of 10,000 synthetic personas covering eight countries: Spain, Mexico, USA, UK, Germany, France, Italy, and Brazil, with 1,250 samples per country. Each persona includes 27 structured fields such as ID, name, country, language, age (18–85), gender, region, city size, education, current and previous occupation, income level (low/middle/high based on country-specific distribution), family type, marital status, number and ages of children, housing type, employment status, living with parents, daily routine, biography, speaking style, hobbies and interests (5), skills (3–6), future goals (1–3), consumer characteristics (shopping style, price sensitivity, brand loyalty, sustainability concern, tech adoption), and personality traits (Big Five: openness, conscientiousness, extraversion, agreeableness, neuroticism; and adult attachment anxiety and avoidance, all 0–1 continuous). All text fields are in the personas country language (English, Spanish, German, French, Italian, Portuguese). The dataset is built on real demographic distributions for age, region, education, income, living arrangements, family structure (and occupation for five countries); names come from real name frequencies; occupations use local names; income reflects household income by education and age for each country. Version 2 introduces realistic occupations (for Spain, France, USA, Mexico, Brazil) and realistic employment rates. It is suitable for testing chatbots and agents, fairness and bias checks, seeding synthetic data, pre-testing surveys and products for specific segments, and role-playing. The dataset is released as JSON Lines files (one per country) under CC BY 4.0, requiring attribution to StrataSynth.
StrataSynth Personas (8 Countries) 数据集概述
基本信息
- 数据集地址:https://huggingface.co/datasets/StrataSynth/stratasynth-personas-8-countries
- 许可证:CC BY 4.0(使用时请注明 StrataSynth)
- 语言:英语(en)、西班牙语(es)、德语(de)、法语(fr)、意大利语(it)、葡萄牙语(pt)
- 任务类别:文本生成(text-generation)
- 数据规模:10K<n<100K
- 标签:synthetic、synthetic-data、personas、synthetic-personas、persona、synthetic-users、stratasynth、user-simulation、agent-evaluation、multilingual、europe、latin-america、psychegraph
数据集内容
- 10,000 个合成人物,来自 8 个国家:西班牙、墨西哥、美国、英国、德国、法国、意大利、巴西。
- 每个国家 1,250 人,每人 27 个字段,共 6 种语言。
- 每个人物包含:身份信息(年龄、地区、工作、家庭、住房)、日常生活作息、说话风格、兴趣爱好、技能、目标,以及数值化的人格档案。传记、日常作息和爱好均基于该人格档案生成。
v2 版本新增内容
- 真实职业:西班牙、法国、美国、墨西哥、巴西的职业遵循各国不同年龄、性别、教育程度和收入人群的实际职业分布,且使用当地称呼(如巴西的 "operadora de caixa"、墨西哥的 "mesero"),每国数百种不同职业。
- 真实收入:
income_level遵循各国按教育和年龄划分的家庭收入实际分布(low= 后 30%,medium= 中间 40%,high= 前 30%),职业与收入保持一致。 - 真实就业:就业与失业比例遵循各国按性别和年龄划分的官方数据。
数据集特点
- 真实人群而非刻板印象:年龄、地区、教育、收入、居住安排和家庭结构(以及 8 国中 5 国的职业)遵循各国真实成年人口。子女年龄与父母年龄匹配。姓名遵循各国按世代划分的真实姓名频率。
- 覆盖欧洲及西语、葡语世界:6 个欧洲和美洲国家使用各自语言,包含地区细节(州、联邦州、大区、自治区)。
- 可筛选的人格:每个人均有大五人格(0-1)和成人依恋(焦虑、回避)。更外向的人有更多社交类爱好;更开放的人有更多文化和学习导向的爱好。
- 可直接作为用户模拟:传记加说话风格正是用户模拟器所需。
- 每个人各不相同:没有两个人共享相同的传记;重复全名比例远低于 1%,与真实人口相似。
主要用途
- 在正式上线前用数千名真实感用户测试聊天机器人和智能体(客户服务、银行、医疗、公共服务)。
- 公平性与偏见检查:向仅年龄、国家或性别不同的人提出相同问题并比较回答。
- 为合成数据播种:"写出这个人会问的问题"能产生比单一提示更多样的指令。
- 作为合成受访者,按细分群体预测试调查、产品和信息。
- 角色扮演和角色创作,从人格条件模型的训练数据到游戏角色。
示例记录
json { "id": "gb-00026", "name": "David Shaw", "country": "GB", "language": "en", "age": 43, "gender": "male", "region": "London", "city_size": "metro", "education": "vocational", "occupation": "telecom installer", "last_occupation": null, "income_level": "medium", "household_type": "family_teen", "marital_status": "partnered", "children": 2, "children_ages": [10, 13], "housing": "rented", "employment_status": "employed", "lives_with_parents": false, "daily_routine": "…", "persona": "…", "speaking_style": "…", "hobbies_and_interests": ["Watching football with friends", "Sunday league five-a-side", "Pub quiz nights", "Family day trips", "DIY around the flat"], "skills": ["Cable installation and testing", "Reading building plans", "Troubleshooting network faults", "Making people laugh", "Basic home repairs"], "goals": ["…"], "consumer": {"shopping_style": "deliberate", "price_sensitivity": "medium", "brand_loyalty": "medium", "sustainability_concern": "medium", "tech_adoption": "mainstream"}, "personality": { "big_five": {"O": 0.5, "C": 0.68, "E": 0.78, "A": 0.21, "N": 0.65}, "attachment": {"anxiety": 0.24, "avoidance": 0.29} } }
字段说明
| 字段 | 类型 | 描述 |
|---|---|---|
id |
string | 国家代码 + 编号,如 es-00001 |
name |
string | 全名。西班牙、墨西哥和巴西按当地习俗使用两个姓氏 |
country |
string | ES、MX、US、GB、DE、FR、IT、BR |
language |
string | 所有文本字段的语言:es、en、de、fr、it、pt |
age |
int | 18 至 85 |
gender |
string | male、female、non_binary |
region |
string | 国内地区(巴西:宏观区域) |
city_size |
string | rural、small、medium、large、metro |
education |
string | secondary、vocational、university、postgrad |
occupation |
string | null |
last_occupation |
string | null |
income_level |
string | low、medium、high,相对于该国 |
household_type |
string | single、couple、family_young、family_teen、empty_nester |
marital_status |
string | single、partnered、married、divorced、widowed |
children |
int | 子女人数 |
children_ages |
list[int] | 子女年龄 |
housing |
string | owned、rented、family_home |
employment_status |
string | employed、unemployed、student、retired、homemaker、other |
lives_with_parents |
bool | null |
daily_routine |
string | 典型的一天,使用人物语言 |
persona |
string | 传记,使用人物语言 |
speaking_style |
string | 说话方式:语气、长度、词汇、习惯,使用人物语言 |
hobbies_and_interests |
list[string] | 五项爱好和兴趣 |
skills |
list[string] | 三到六项技能 |
goals |
list[string] | 未来一到三个目标 |
consumer |
object | shopping_style、price_sensitivity、brand_loyalty、sustainability_concern、tech_adoption |
personality.big_five |
object | O 开放性、C 尽责性、E 外向性、A 宜人性、N 神经质,0 至 1 |
personality.attachment |
object | 成人依恋 anxiety 和 avoidance,0 至 1 |
数据一览
- 10,000 人,每国 1,250 人;年龄 18-85 岁(中位数 47)。
- 就业:59% 在职,19% 退休,9% 居家,4% 失业,5% 在读。
- 职业:西班牙、法国、美国、墨西哥和巴西每国 243 至 401 种不同职业。
- 住房:48% 拥有住房,38% 租房,14% 住在家庭住宅。
- 文本:传记约 145 词;每个人都有日常作息、说话风格、爱好、技能和目标。
配置与文件
- 配置名称:
default - 每个国家一个 JSON Lines 文件,每国为一个 split:
split: es→personas_es.jsonlsplit: mx→personas_mx.jsonlsplit: us→personas_us.jsonlsplit: gb→personas_gb.jsonlsplit: de→personas_de.jsonlsplit: fr→personas_fr.jsonlsplit: it→personas_it.jsonlsplit: br→personas_br.jsonl
使用代码
python from datasets import load_dataset, concatenate_datasets ds = load_dataset("StrataSynth/stratasynth-personas-8-countries") # one split per country es = ds["es"] # 1,250 people from Spain everyone = concatenate_datasets(list(ds.values())) # all 10,000
单国数据集
每个国家也单独发布,并配有各自的图表:
- 西班牙:https://huggingface.co/datasets/StrataSynth/stratasynth-personas-spain
- 墨西哥:https://huggingface.co/datasets/StrataSynth/stratasynth-personas-mexico
- 美国:https://huggingface.co/datasets/StrataSynth/stratasynth-personas-united-states
- 英国:https://huggingface.co/datasets/StrataSynth/stratasynth-personas-united-kingdom
- 德国:https://huggingface.co/datasets/StrataSynth/stratasynth-personas-germany
- 法国:https://huggingface.co/datasets/StrataSynth/stratasynth-personas-france
- 意大利:https://huggingface.co/datasets/StrataSynth/stratasynth-personas-italy
- 巴西:https://huggingface.co/datasets/StrataSynth/stratasynth-personas-brazil
以上全部归属于 StrataSynth Personas 集合:https://huggingface.co/collections/StrataSynth/stratasynth-personas-6ab96c1be53faa6c4c596031
联系方式
这些人物来源于 StrataSynth 的 Humans Engine。如需特定细分群体或国家的人物,或希望对其进行访谈,请联系:hello@stratasynth.com
注意事项
- 每个人物完全为合成数据,并非真实人物或任何机构的记录,与真实人物的任何相似均为巧合。
- 姓名取自各国真实姓名频率,因此合成人物可能与真实人物同名。





