stratasynth-personas-brazil
收藏资源简介:
StrataSynth Personas Brazil 是一个包含 1,250 个合成人物画像的数据集,全部以巴西葡萄牙语书写。每个人物都包含详细的传记、说话风格、日常安排、爱好、技能、目标以及基于大五人格和成人依恋模型的人格测量。数据集共有 27 个字段,涵盖年龄、性别、地区、教育、职业、收入、家庭结构、居住状况等人口统计信息。该数据集基于巴西真实人口统计数据生成,确保年龄、地区、教育、职业、收入和家庭结构符合实际分布。v2 版本引入了真实职业(以当地语言命名)、与教育和年龄相匹配的收入分布以及符合官方统计的就业率。数据可广泛应用于测试面向巴西用户的聊天机器人和智能体、进行年龄/性别/区域的公平性检查、生成葡萄牙语合成数据(如模拟用户提问)、以及用于调查和产品的预测试。所有人物均为合成,不涉及真实个人,名字基于真实姓名频率。
StrataSynth Personas Brazil is a dataset containing 1,250 synthetic personas, all written in Brazilian Portuguese. Each persona includes detailed biography, speaking style, daily schedule, hobbies, skills, goals, and personality measurements based on the Big Five and Adult Attachment models. The dataset has a total of 27 fields covering demographic information such as age, gender, region, education, occupation, income, family structure, and housing status. It is generated based on real Brazilian demographic statistics, ensuring that age, region, education, occupation, income, and family structure follow actual distributions. Version v2 introduces real occupations (named in the local language), income distribution matched to education and age, and employment rates consistent with official statistics. The data can be widely used for testing chatbots and agents targeting Brazilian users, fairness checks by age/gender/region, generating Portuguese synthetic data (e.g., simulating user questions), and pre-testing surveys and products. All personas are synthetic and do not involve real individuals; names are based on real name frequencies.
StrataSynth Personas Brazil 数据集概述
基本信息
- 数据集名称:StrataSynth Personas Brazil
- 数据集地址:https://huggingface.co/datasets/StrataSynth/stratasynth-personas-brazil
- 许可证:CC BY 4.0
- 语言:葡萄牙语(巴西)
- 数据规模:1,250 条,规模类别为 1K<n<10K
- 任务类别:文本生成
- 数据文件:personas_br.jsonl(train 划分)
- 数据标签:synthetic、synthetic-data、personas、synthetic-personas、persona、synthetic-users、user-simulation、agent-evaluation、brazil、stratasynth
数据集简介
该数据集包含 1,250 个来自巴西的合成人物,使用葡萄牙语(巴西)撰写。每个合成人物包含传记、说话方式、爱好、技能、目标以及测量的人格特征。
- 1,250 个人物
- 每人 27 个字段
- 覆盖 5 个宏观区域
该数据集是 StrataSynth Personas 集合的一部分(八个国家、10,000 人)。完整集合地址:https://huggingface.co/datasets/StrataSynth/stratasynth-personas-8-countries。
v2 版本新增内容
- 真实职业:工作内容遵循每个年龄段、性别、教育程度和收入人群在当地实际从事的工作,职业名称按当地说法命名(每个国家有数百种不同职业,例如巴西的 "operadora de caixa")。
- 真实收入:
income_level遵循各国按教育程度和年龄划分的真实家庭收入分布(low= 后 30%,medium= 中间 40%,high= 前 30%),且职业与之保持一致。 - 真实就业:就业与失业比例遵循各国按性别和年龄划分的官方比率。
数据集特点
- 真实人口,而非刻板印象:年龄、地区、教育、职业、收入、居住安排和家庭结构遵循该国真实成年人口。子女年龄与父母年龄匹配,姓名遵循各代际的真实姓名频率。
- 以葡萄牙语(巴西)原生撰写:传记、日常作息、说话风格、爱好、技能和目标均为母语撰写,而非翻译。
- 可筛选的人格:为每个人物提供大五人格(0-1)和成人依恋。更外向的人拥有更多社交类爱好;更开放的人拥有更多文化和学习导向的爱好。
- 可直接作为用户使用:传记加说话风格正是用户模拟器所需的:将其放入提示词中,即可得到一个来自巴西的顾客、患者或受访者。
人口特征
- 年龄:18-85 岁,中位数 42 岁。
- 就业:62% 就业,12% 退休,20% 居家,4% 失业,2% 在读。
- 住房:50% 自有住房,45% 租房,5% 住在家庭住房。
主要用途
- 在正式上线前,用真实的本地用户测试为巴西构建的聊天机器人和智能体。
- 跨年龄、性别和地区进行公平性与偏见检查。
- 为葡萄牙语(巴西)合成数据做种子:"写出这个人会问的问题"。
- 作为合成受访者,预测试调查、产品和信息传达。
字段说明
| 字段 | 类型 | 描述 |
|---|---|---|
id |
string | 国家代码 + 编号,例如 es-00001 |
name |
string | 全名。按当地习俗,西班牙、墨西哥和巴西使用两个姓氏 |
country |
string | ES、MX、US、GB、DE、FR、IT、BR |
language |
string | 所有文本字段的语言:es、en、de、fr、it、pt |
age |
int | 18 至 85 |
gender |
string | male、female、non_binary |
region |
string | 国家内的区域(巴西:宏观区域) |
city_size |
string | rural、small、medium、large、metro |
education |
string | secondary、vocational、university、postgrad |
occupation |
string | null |
last_occupation |
string | null |
income_level |
string | low、medium、high,相对于该国 |
household_type |
string | single、couple、family_young、family_teen、empty_nester |
marital_status |
string | single、partnered、married、divorced、widowed |
children |
int | 子女数量 |
children_ages |
list[int] | 子女年龄 |
housing |
string | owned、rented、family_home |
employment_status |
string | employed、unemployed、student、retired、homemaker、other |
lives_with_parents |
bool | null |
daily_routine |
string | 典型的一天,使用人物语言 |
persona |
string | 传记,使用人物语言 |
speaking_style |
string | 说话方式:语气、长度、词汇、习惯,使用人物语言 |
hobbies_and_interests |
list[string] | 五项爱好和兴趣 |
skills |
list[string] | 三至六项技能 |
goals |
list[string] | 未来一到三个目标 |
consumer |
object | shopping_style、price_sensitivity、brand_loyalty、sustainability_concern、tech_adoption |
personality.big_five |
object | O 开放性、C 尽责性、E 外向性、A 宜人性、N 神经质,0 至 1 |
personality.attachment |
object | 成人依恋 anxiety 和 avoidance,0 至 1 |
加载方式
python from datasets import load_dataset people = load_dataset("StrataSynth/stratasynth-personas-brazil", split="train")
备注
- 每个人物均为完全合成。他们不是真实人物或任何机构的记录,与真实人物的任何相似之处均为巧合。
- 姓名取自真实姓名频率,因此合成人物可能与真实人物同名。
- 使用时请注明来源 StrataSynth。





