Xalitobeirut/Nemotron-Personas-Korea
收藏资源简介:
Nemotron-Personas-Korea是一个基于韩国真实世界人口统计、地理和性格特征分布合成的开源人物角色数据集(CC BY 4.0)。该数据集旨在广泛反映韩国人口的多样性和特征,包含姓名、性别、年龄、婚姻状况、教育水平、职业和居住地区等属性。数据集使用NVIDIA的NeMo Data Designer企业级合成数据生成复合AI系统创建,结合了专有的概率图模型和Apache-2.0许可的google/gemma-4-31B-it模型。数据集包含100万条记录,涵盖700万个人物角色,26个字段,包括7个人物角色字段、6个人物角色属性字段、12个人口统计和地理背景字段以及1个唯一标识符。该数据集支持韩国模型开发者构建包含重要地区特定人口统计和文化背景的主权AI系统,可用于扩大合成数据的多样性、减轻数据和模型偏见以及提高模型响应的多样性。数据集在CC BY 4.0许可下免费供商业和非商业使用。
Nemotron-Personas-Korea is an open-source persona dataset (CC BY 4.0) synthesized based on real-world demographic, geographic, and personality trait distributions of South Korea. It is designed to broadly reflect the diversity and characteristics of the South Korean population, including attributes such as name, sex, age, marital status, education level, occupation, and region of residence. The dataset was created using NVIDIAs NeMo Data Designer, an enterprise-grade compound AI system for synthetic data generation, leveraging a proprietary probabilistic graphical model and the Apache-2.0 licensed google/gemma-4-31B-it model. It contains 7M personas across 1M records with 26 fields: 7 persona fields, 6 persona attribute fields, 12 demographic & geographic contextual fields, and 1 unique identifier. This dataset supports South Korean model builders in developing Sovereign AI systems that incorporate important region-specific demographics and cultural context. It can be used to expand the diversity of synthetic data, mitigate data and model bias, and improve the diversity of model responses. The dataset is freely available for both commercial and non-commercial use under the CC BY 4.0 license.




