JayElise/Nemotron-Personas-Korea
收藏资源简介:
Nemotron-Personas-Korea是一个基于韩国真实世界人口统计、地理和性格特征分布合成的开源人物角色数据集(CC BY 4.0),旨在广泛反映韩国人口的多样性和特征。作为首个大规模韩语人物角色数据集,它包括姓名、性别、年龄、婚姻状况、教育水平、职业和居住地区等属性,这些属性基于韩国统计信息服务(KOSIS)、韩国最高法院、国民健康保险公团、韩国农村经济研究院和NAVER Cloud的官方统计资料合成。该数据集支持韩国模型开发者构建包含重要地区特定人口统计和文化背景的主权AI系统,可用于扩大主权AI模型开发的合成数据多样性、缓解数据和模型偏见,并提高模型响应的多样性。数据集使用NeMo Data Designer(企业级合成数据生成复合AI系统)创建,利用专有的概率图模型、Apache-2.0许可的google/gemma-4-31B-it模型以及Data Designer中的验证和评估方法。数据集包含100万条记录、26个字段(7个人物角色字段、6个人物角色属性字段、12个人口统计和地理上下文字段、1个唯一标识符),覆盖17个省份和252个区,包含209,000个唯一姓名和7种人物角色类型(职业、体育、艺术、旅行、烹饪、家庭、简洁)。数据集仅包含成年(19岁及以上)人物角色,完全基于合成数据,但反映真实分布。
Nemotron-Personas-Korea is an open-source synthetic persona dataset licensed under CC BY 4.0, synthesized based on the real-world demographic, geographic, and personality trait distributions of the Korean population, aiming to broadly reflect the diversity and distinctive features of the Korean populace. As the first large-scale Korean persona dataset, it encompasses attributes including name, gender, age, marital status, education level, occupation, and residential region. These attributes are generated using official statistical materials from the Korea Statistical Information Service (KOSIS), the Supreme Court of Korea, the National Health Insurance Service, the Korea Rural Economic Institute, and NAVER Cloud. This dataset enables Korean model developers to build sovereign AI systems incorporating critical region-specific demographic and cultural backgrounds, and can be used to expand the diversity of synthetic data for sovereign AI model development, mitigate biases in both data and models, and enhance the diversity of model responses. The dataset was created using NeMo Data Designer, an enterprise-grade composite AI system for synthetic data generation, which leverages proprietary probabilistic graphical models, the Apache-2.0-licensed google/gemma-4-31B-it model, as well as validation and evaluation methods integrated within Data Designer. It contains 1 million records and 26 fields, specifically 7 persona-related fields, 6 persona attribute fields, 12 demographic and geographic context fields, and 1 unique identifier. The dataset covers 17 provinces and 252 administrative districts, includes 209,000 unique names, and features 7 persona categories: occupation, sports, art, travel, cooking, family, and concise. All included personas are adults aged 19 years and above; the dataset is fully synthetic, yet accurately reflects real-world population distributions.



