Reza2kn/uncgpt-personas
收藏资源简介:
UncGPT — Persona Pool是一个包含1,000,069个合成角色和15,000个名称注册表的数据集,作为所有UncGPT照顾对话生成的基础。它是多语言的,支持英语、西班牙语、中文、印地语、波斯语、斯瓦希里语、法语、葡萄牙语、孟加拉语、约鲁巴语、他加禄语,主要用于文本生成和特征提取任务。数据集包括角色字段(如标识符、性别、语言、文化、年龄带、地理信息等)和JSON负载(包含施瓦茨价值观、HEXACO特质、家庭/社交配置、就业、家庭状况、应对风格等)。所有数据均为合成生成,无人类数据抓取,基于公开许可的种子词典生成。
UncGPT — Persona Pool is a dataset containing 1,000,069 synthetic personas plus a 15,000-name registry, serving as the substrate for all UncGPT caregiving conversation generation. It is multilingual, supporting English, Spanish, Chinese, Hindi, Persian, Swahili, French, Portuguese, Bengali, Yoruba, and Tagalog, and is designed for text-generation and feature-extraction tasks. The dataset includes persona fields (such as identifier, gender, language, culture, age band, geographic information) and a serialized JSON payload (covering Schwartz values, HEXACO traits, family/social configuration, employment, household, coping style, etc.). All personas are synthetically generated with no human data scraped, based on publicly licensed seed lexicons.



