tantara/Nemotron-Personas-Japan-Qwen3-0.6B-embedding
收藏资源简介:
该数据集名为Nemotron-Personas-Japan-Qwen3-0.6B-embedding,是一个嵌入向量数据集,专门为源数据集nvidia/Nemotron-Personas-Japan(日本人物角色数据集)使用Qwen/Qwen3-Embedding-0.6B模型计算得到。嵌入维度为1024,包含100万行数据(源自源数据集700万行的前100万行)。嵌入的列涵盖了19个字段:专业人物角色、体育人物角色、艺术人物角色、旅行人物角色、烹饪人物角色、人物角色、文化背景、技能与专长、爱好与兴趣、职业目标与抱负、性别、年龄、婚姻状况、教育水平、职业、地区、区域、都道府县和国家。数据集的schema包括两列:uuid(字符串类型,主键,用于连接回源数据集)和embedding(float32列表类型,表示1024维的嵌入向量)。该数据集可用于自然语言处理任务,如人物角色分析、嵌入向量检索等,并遵循cc-by-4.0许可证。
The dataset Nemotron-Personas-Japan-Qwen3-0.6B-embedding is an embeddings dataset computed for the source dataset nvidia/Nemotron-Personas-Japan using the Qwen/Qwen3-Embedding-0.6B model. It has an embedding dimension of 1024 and contains 1,000,000 rows (the first 1M rows of the 7M-row source dataset). The columns embedded include 19 fields: professional_persona, sports_persona, arts_persona, travel_persona, culinary_persona, persona, cultural_background, skills_and_expertise, hobbies_and_interests, career_goals_and_ambitions, sex, age, marital_status, education_level, occupation, region, area, prefecture, and country. The schema consists of two columns: uuid (string type, primary key for joining back to the source dataset) and embedding (list of float32, representing a 1024-dimensional embedding vector). This dataset is suitable for NLP tasks such as persona analysis and embedding retrieval, and is licensed under cc-by-4.0.



