dataj12/Nemotron-Personas-Korea
收藏资源简介:
Nemotron-Personas-Korea是一个基于韩国真实世界人口统计、地理和人格特征分布合成的开源人物数据集(CC BY 4.0),设计用于广泛反映韩国人口的多样性和特征。作为首个大规模韩语人物数据集,它包含姓名、性别、年龄、婚姻状况、教育水平、职业和居住地区等属性,这些属性基于韩国统计信息院(KOSIS)、最高法院、国民健康保险公团、农村经济研究院和NAVER Cloud的官方统计数据合成。数据集旨在支持韩国模型开发者构建包含重要地区特定人口统计和文化背景的主权AI系统,可用于扩展主权AI模型开发的合成数据多样性、减轻数据和模型偏见,并提高模型响应多样性。特别地,与现有人物数据集相比,它更忠实地反映了多个维度上的真实人口分布,包括年龄(如老年人群)、地区(如农村地区)、教育水平和职业。数据集使用NeMo Data Designer(一个企业级合成数据生成复合AI系统)创建,利用专有的概率图模型、Apache-2.0许可的google/gemma-4-31B-it模型以及Data Designer中包含的验证和评估方法。数据集包含100万条记录、26个字段(7个人物字段、6个人物属性字段、12个人口统计和地理上下文字段以及1个唯一标识符),覆盖17个省份和252个地区,包含209,000个唯一名称(118个姓氏,21,400个名字)和7种人物类型(专业、体育、艺术、旅行、烹饪、家庭、简洁),以及额外的自然语言人物属性(文化背景、技能与专业知识、职业目标与抱负、爱好与兴趣)。数据集仅包含韩国法律规定的成年年龄(19岁及以上)的人物,并假设变量间独立性,未建模因素交互作用。数据完全人工合成,任何与真实人物的相似性纯属巧合。
Nemotron-Personas-Korea is an open-source persona dataset (CC BY 4.0) synthesized based on real-world demographic, geographic, and personality trait distributions of South Korea, designed to broadly reflect the diversity and characteristics of the South Korean population. As the first large-scale Korean-language persona dataset, it includes attributes such as name, sex, age, marital status, education level, occupation, and region of residence, all synthesized using official statistics from the Korean Statistical Information Service (KOSIS), the Supreme Court of Korea, the National Health Insurance Service, the Korea Rural Economic Institute, and NAVER Cloud. The dataset supports South Korean model builders in developing Sovereign AI systems that incorporate important region-specific demographics and cultural context, and can be used to expand the diversity of synthetic data for sovereign AI model development, mitigate data and model bias, and improve the diversity of model responses. In particular, compared to existing persona datasets, it is designed to more faithfully reflect real population distributions across multiple dimensions, including age (e.g., elderly populations), region (e.g., rural areas), education level, and occupation. The dataset was created using NeMo Data Designer, an enterprise-grade compound AI system for synthetic data generation, leveraging a proprietary probabilistic graphical model (PGM), the Apache-2.0 licensed google/gemma-4-31B-it model, and the validation and evaluation methods included in Data Designer. It contains 1 million records with 26 fields (7 persona fields, 6 persona attribute fields, 12 demographic & geographic contextual fields, and 1 unique identifier), comprehensive geographic coverage across 17 provinces and 252 districts, 209,000 unique names (118 surnames, 21,400 given names), 7 persona types (professional, sports, arts, travel, culinary, family, concise), and additional natural language persona attributes (cultural background, skills & expertise, career goals & ambitions, hobbies & interests). The dataset includes only personas of adult age (19 years and older by South Korean law), applies independence assumptions between variables without modeling interaction effects, and is completely artificially generated with any similarity to actual persons being purely coincidental.



