ilikesleep/Nemotron-Personas-Korea
收藏资源简介:
Nemotron-Personas-Korea是一个基于韩国真实世界人口统计、地理和个性特征分布合成的开源人物角色数据集(CC BY 4.0),旨在广泛反映韩国人口的多样性和特征。作为首个大规模韩语人物角色数据集,它包含姓名、性别、年龄、婚姻状况、教育水平、职业、居住地区等属性,这些属性基于韩国统计信息服务(KOSIS)、韩国最高法院、国民健康保险公团、韩国农村经济研究院和NAVER Cloud的官方统计数据合成。数据集支持韩国模型开发者构建融入重要区域特定人口统计和文化背景的Sovereign AI系统。它可用于扩大主权AI模型开发的合成数据多样性、缓解数据和模型偏差,并提高模型响应的多样性。与现有人物角色数据集相比,它更忠实地反映了真实人口分布在多个维度上的情况,包括年龄(例如老年人口)、地区(例如农村地区)、教育水平和职业。数据集使用NeMo Data Designer(一个企业级复合AI系统,用于合成数据生成)创建,利用了专有的概率图模型、Apache-2.0许可的google/gemma-4-31B-it模型以及Data Designer中包含的验证和评估方法。数据集包含700万个人物角色,分布在100万条记录中,有26个字段:7个人物角色字段、6个人物角色属性字段、12个人口统计和地理上下文字段以及1个唯一标识符。它涵盖17个省份和252个地区的全面地理覆盖,包含209,000个唯一姓名(118个姓氏,21,400个名字),以及7种人物角色类型:职业、体育、艺术、旅行、烹饪、家庭和简洁类型,还包括额外的自然语言人物角色属性,如文化背景、技能与专业知识、职业目标与抱负、爱好与兴趣。数据集完全免费用于商业和非商业用途。
Nemotron-Personas-Korea is an open-source synthetic persona dataset licensed under CC BY 4.0, built based on the real-world demographic, geographic, and personality trait distributions of the Korean population, aiming to broadly reflect the diversity and characteristics of the Korean populace. As the first large-scale Korean persona dataset, it encompasses attributes such as name, gender, age, marital status, educational attainment, occupation, and residential area. These attributes are synthesized using official statistical data from the Korea Statistical Information Service (KOSIS), the Supreme Court of Korea, the National Health Insurance Service, the Korea Rural Economic Institute, and NAVER Cloud. The dataset enables Korean model developers to construct Sovereign AI systems that integrate critical region-specific demographic and cultural contexts. It can be utilized to expand the diversity of synthetic data for sovereign AI model development, mitigate data and model biases, and improve the diversity of model responses. Compared to existing persona datasets, it more faithfully mirrors real population distributions across multiple dimensions, including age (e.g., elderly population), region (e.g., rural areas), educational attainment, and occupation. The dataset was created using NeMo Data Designer, an enterprise-grade composite AI system for synthetic data generation. It leverages proprietary probabilistic graphical models, the Apache-2.0 licensed google/gemma-4-31B-it model, and the validation and evaluation methods included in Data Designer. It contains 7 million personas distributed across 1 million records, with 26 fields in total: 7 persona fields, 6 persona attribute fields, 12 demographic and geographic context fields, and 1 unique identifier. It features comprehensive geographic coverage across 17 provinces and 252 regions, includes 209,000 unique names (118 surnames and 21,400 given names), and covers 7 persona categories: occupation, sports, art, travel, cooking, family, and concise profile. It also includes additional natural language persona attributes such as cultural background, skills and expertise, career goals and aspirations, and hobbies and interests. The dataset is fully free for both commercial and non-commercial use.



