nguyenthuyan/Nemotron-Personas-Korea
收藏资源简介:
Nemotron-Personas-Korea是基于韩国真实人口统计、地理和个性特征分布合成的开源人物角色数据集(CC BY 4.0),旨在广泛反映韩国人口的多样性和特性。作为首个大规模韩语人物角色数据集,它包含姓名、性别、年龄、婚姻状况、教育水平、职业、居住地区等属性,这些属性基于韩国统计信息服务(KOSIS)、大法院、国民健康保险公团、农村经济研究院和NAVER Cloud的官方统计数据合成。该数据集支持韩国模型开发者构建融入重要地区特定人口统计和文化背景的主权AI系统,可用于扩展主权AI模型开发的合成数据多样性、减轻数据和模型偏见、提高模型响应多样性。与现有人物角色数据集相比,它更忠实地反映了多个维度(如年龄、地区、教育水平、职业)的真实人口分布。数据集使用NeMo Data Designer(企业级合成数据生成复合AI系统)创建,包含700万个人物角色、26个字段(7个人物角色字段、6个人物角色属性字段、12个人口统计和地理上下文字段、1个唯一标识符),覆盖17个省份和252个地区,以及209,000个唯一姓名。数据集可自由用于商业和非商业用途。
Nemotron-Personas-Korea is an open-source synthetic persona dataset licensed under CC BY 4.0, developed based on the real demographic, geographic, and personality trait distributions of South Korea, aiming to comprehensively reflect the diversity and characteristics of the Korean population. As the first large-scale Korean persona dataset, it includes attributes such as name, gender, age, marital status, education level, occupation, and residential region. These attributes are synthesized using official statistical data from the Korean Statistical Information Service (KOSIS), the Supreme Court, National Health Insurance Service, Rural Economic Institute, and NAVER Cloud. This dataset enables Korean model developers to build sovereign AI systems that incorporate critical region-specific demographic and cultural contexts, and can be used to expand the diversity of synthetic data for sovereign AI model development, mitigate data and model biases, and enhance the diversity of model responses. Compared to existing persona datasets, it more faithfully reflects real population distributions across multiple dimensions including age, region, education level, and occupation. The dataset was created using NeMo Data Designer, an enterprise-grade composite AI system for synthetic data generation, and contains 7 million personas across 26 fields: 7 persona fields, 6 persona attribute fields, 12 demographic and geographic context fields, and 1 unique identifier. It covers 17 provinces and 252 regions, with 209,000 unique names. The dataset is freely available for both commercial and non-commercial use.



