nmlab2025/Nemotron-Personas-Korea
收藏资源简介:
Nemotron-Personas-Korea是一个基于韩国真实世界人口统计、地理和人格特征分布的开源合成人物数据集(CC BY 4.0)。它旨在广泛反映韩国人口的多样性和特征。作为首个大规模的韩语人物数据集,它包含了姓名、性别、年龄、婚姻状况、教育水平、职业和居住地区等属性,这些属性均基于韩国统计信息服务(KOSIS)、韩国最高法院、国民健康保险公团、韩国农村经济研究院和NAVER Cloud的官方统计数据合成。该数据集支持韩国模型开发者开发包含重要地区特定人口统计和文化背景的Sovereign AI系统。它可用于扩大主权AI模型开发的合成数据多样性,缓解数据和模型偏见,并提高模型响应的多样性。数据集使用NVIDIA的NeMo Data Designer创建,这是一个企业级复合AI系统,用于合成数据生成。数据集包含700万个人物,分布在100万条记录中,涵盖26个字段,包括7个人物类型字段、6个人物属性字段、12个人口统计和地理上下文字段以及1个唯一标识符。数据集在CC BY 4.0许可下免费供商业和非商业使用。
Nemotron-Personas-Korea is an open-source synthetic persona dataset licensed under CC BY 4.0, built based on the distribution of real-world demographic, geographic, and personality traits of the Korean population. It is designed to comprehensively reflect the diversity and characteristics of the Korean populace. As the first large-scale Korean persona dataset, it encompasses attributes including name, gender, age, marital status, educational attainment, occupation, and residential region, all synthesized from official statistical data provided by the Korea Statistical Information Service (KOSIS), Supreme Court of Korea, National Health Insurance Service, Korea Rural Economic Institute, and NAVER Cloud. This dataset enables Korean model developers to construct Sovereign AI systems that incorporate critical region-specific demographic and cultural backgrounds. It can be utilized to augment the diversity of synthetic data for sovereign AI model development, alleviate data and model biases, and improve the diversity of model responses. The dataset was developed using NVIDIA’s NeMo Data Designer, an enterprise-grade composite AI system tailored for synthetic data generation. It contains 7 million personas spread across 1 million records, covering 26 fields in total: 7 persona type fields, 6 persona attribute fields, 12 demographic and geographic context fields, and 1 unique identifier. The dataset is freely accessible for both commercial and non-commercial purposes under the CC BY 4.0 license.



