遇见数据集

charaybids/Nemotron-Personas-Singapore

收藏
Hugging Face2026-05-13 更新2026-05-31 收录
官方服务:

资源简介:

Nemotron-Personas-Singapore是一个开源(CC BY 4.0)的合成生成人物数据集,基于新加坡的真实世界人口统计、地理和人格特征分布构建,以捕捉新加坡人口的多样性和丰富性。该数据集是Nemotron-Personas-USA的变体,是首个与新加坡的姓名、性别、年龄、民族、宗教、婚姻状况和职业等属性统计数据对齐的新加坡数据集。数据集包含888k个人物描述,分布在148k条记录中,总计118M个令牌(包括48M个人物令牌)。数据包含20个字段,包括6个人物字段(如专业人物、体育人物、艺术人物等)和14个上下文字段(如文化背景、技能、年龄、职业等)。数据集支持新加坡模型开发者开发主权AI系统,融入区域特定的人口统计和文化背景,提高合成生成数据的多样性,减轻偏见,防止模型崩溃。数据通过NeMo Data Designer生成,使用概率图模型和GPT-OSS-120B模型。数据集仅包含成年人人物,基于2024年新加坡人口普查数据构建。

Nemotron-Personas-Singapore is an open-source (CC BY 4.0) dataset of synthetically-generated personas grounded in real-world demographic, geographic and personality trait distributions in Singapore to capture the diversity and richness of the Singaporean population. It is a variant of Nemotron-Personas-USA and the first Singaporean dataset of its kind aligned with statistics for names, sex, age, ethnicity, religion, marital status and occupation among other attributes. The dataset contains 888k personas across 148k records, with 118M tokens including 48M persona tokens. It includes 20 fields comprising 6 persona fields (e.g., professional_persona, sports_persona) and 14 contextual fields (e.g., cultural_background, skills_and_expertise, age, occupation). The dataset supports Singaporean model builders in developing Sovereign AI systems that incorporate region-specific demographics and cultural context, improving diversity of synthetically-generated data, mitigating biases, and preventing model collapse. It is produced using NeMo Data Designer with a Probabilistic Graphical Model and GPT-OSS-120B model. The dataset focuses on adults only and is based on the 2024 census of Singapore.

提供机构:
charaybids
二维码
社区交流群
二维码
科研交流群
商业服务