liyah1616/Nemotron-Personas-Korea
收藏资源简介:
Nemotron-Personas-Korea是一个基于韩国真实人口统计、地理和人格特质分布合成的开源角色数据集(采用CC BY 4.0许可),旨在广泛反映韩国人口的多样性和特征。作为首个大规模韩语角色数据集,它包含姓名、性别、年龄、婚姻状况、教育水平、职业、居住地区等属性,这些属性基于韩国统计信息服务(KOSIS)、韩国最高法院、国民健康保险公团、韩国农村经济研究院和NAVER Cloud的官方统计数据合成。该数据集支持韩国模型开发者构建融入重要地区特定人口统计和文化背景的“主权AI”系统,可用于扩展主权AI模型开发的合成数据多样性、缓解数据和模型偏见,并提高模型响应的多样性。与现有角色数据集相比,它更忠实地反映了真实人口分布在多个维度上的情况,包括年龄(如老年人口)、地区(如农村)、教育水平和职业。数据集使用NeMo Data Designer(一个企业级复合AI系统)创建,利用了专有的概率图模型、Apache-2.0许可的google/gemma-4-31B-it模型以及Data Designer中的验证和评估方法。数据集包含100万条记录、700万个角色、26个字段(7个角色字段、6个角色属性字段、12个人口统计和地理上下文字段、1个唯一标识符),覆盖韩国17个省份和252个区,包含20.9万个唯一姓名(118个姓氏,2.14万个名字)和7种角色类型(职业、体育、艺术、旅行、烹饪、家庭、简洁)。数据集仅包含成年(19岁及以上)角色,适用于商业和非商业用途。
Nemotron-Personas-Korea is an open-source synthetic character dataset licensed under CC BY 4.0, synthesized based on the real demographic, geographic, and personality trait distributions of the Korean population, aiming to broadly reflect the diversity and characteristics of the Korean populace. As the first large-scale Korean character dataset, it encompasses attributes including name, gender, age, marital status, educational attainment, occupation, and residential area. These attributes are synthesized using official statistical data from the Korea Statistical Information Service (KOSIS), the Supreme Court of the Republic of Korea, the National Health Insurance Service of the Republic of Korea, the Korea Rural Economic Institute (KREI), and NAVER Cloud. This dataset enables Korean model developers to build "sovereign AI" systems that integrate critical region-specific demographic and cultural backgrounds. It can be used to expand the diversity of synthetic data for sovereign AI model development, mitigate biases in both data and models, and enhance the diversity of model responses. Compared with existing character datasets, it more faithfully mirrors real population distributions across multiple dimensions, including age (e.g., elderly populations), regions (e.g., rural areas), educational attainment, and occupation. The dataset was constructed using NeMo Data Designer, an enterprise-grade composite AI system, leveraging proprietary probabilistic graphical models, the Apache-2.0 licensed google/gemma-4-31B-it model, and validation and evaluation methods integrated within Data Designer. It contains 1 million records, 7 million character profiles, 26 total fields (7 character-related fields, 6 character attribute fields, 12 demographic and geographic context fields, and 1 unique identifier), covers all 17 provinces and 252 districts of the Republic of Korea, includes 209,000 unique names (118 surnames and 21,400 given names), and features 7 character template types: occupation, sports, art, travel, cooking, family, and concise. The dataset only includes adult characters aged 19 years and above, and is applicable for both commercial and non-commercial use.



