ChloeLynn/Nemotron-Personas-Korea
收藏资源简介:
Nemotron-Personas-Korea是一个基于韩国真实人口统计、地理和个性特征分布的开源合成人物数据集(CC BY 4.0)。它旨在广泛反映韩国人口的多样性和特征,是首个大规模韩语人物数据集。数据集包含姓名、性别、年龄、婚姻状况、教育水平、职业和居住地区等属性,这些属性基于韩国统计信息服务中心(KOSIS)、韩国最高法院、国民健康保险公团、农村经济研究院和NAVER Cloud的官方统计数据合成。该数据集支持韩国模型开发者开发包含重要地区特定人口统计和文化背景的主权AI系统。数据集可用于扩大主权AI模型开发的合成数据多样性,缓解数据和模型偏见,并提高模型响应的多样性。数据集由NVIDIA Corporation使用NeMo Data Designer企业级合成数据生成复合AI系统创建,并利用专有的概率图模型(PGM)、Apache-2.0许可的google/gemma-4-31B-it模型以及Data Designer中包含的验证和评估方法。数据集可免费用于商业和非商业用途。
Nemotron-Personas-Korea is an open-source persona dataset (CC BY 4.0) synthesized based on real-world demographic, geographic, and personality trait distributions of South Korea. It is designed to broadly reflect the diversity and characteristics of the South Korean population. As the first large-scale Korean-language persona dataset, it includes attributes such as name, sex, age, marital status, education level, occupation, and region of residence, all synthesized using official statistics from the Korean Statistical Information Service (KOSIS), the Supreme Court of Korea, the National Health Insurance Service, and the Korea Rural Economic Institute, and NAVER Cloud. The dataset supports South Korean model builders in developing Sovereign AI systems that incorporate important region-specific demographics and cultural context. This dataset can be used to expand the diversity of synthetic data for sovereign AI model development, mitigate data and model bias, and improve the diversity of model responses. The dataset was created by NVIDIA Corporation using NeMo Data Designer, an enterprise-grade compound AI system for synthetic data generation, leveraging a proprietary probabilistic graphical model (PGM), the Apache-2.0 licensed google/gemma-4-31B-it model, and the validation and evaluation methods included in Data Designer. The dataset is freely available for both commercial and non-commercial use.




