Ideal-TSUNDERE-Loli-Girl-Japanese-v1
收藏资源简介:
该数据集是一个专为日语聊天模型设计的高质量对话数据集,完全由人工手动创建,未使用AI干预。数据内容聚焦于体现“萝莉”、“撒娇”、“依赖”、“安心感”和“亲密性”等特质的自然对话风格。数据集采用JSONL格式,包含120个独立样本,每个样本由用户(user)和助手(assistant)角色之间的对话消息序列构成,强调短句、高密度交互,并特别重视情感表达、距离感以及营造柔和、亲密的对话氛围。其设计目标是优化对话式微调(conversational fine-tuning),适用于大型语言模型的监督微调、LoRA训练以及日语对话模型和情感对话模型的研究与增强。推荐与Gemma、Qwen、Llama、Mistral、OpenHermes等模型系列配合使用。注意事项指出,由于数据集专注于特定对话风格,在基础模型上可能产生重复、情感过拟合或依赖偏差等问题,建议根据需要与通用对话数据或知识数据混合训练。数据集采用MIT许可证发布。
This dataset is a high-quality dialogue dataset specifically designed for Japanese chat models, entirely manually created without AI intervention. The content focuses on natural conversational styles that embody traits such as "loli", "coquetry", "dependence", "sense of security", and "intimacy". The dataset is in JSONL format, containing 120 independent samples, each consisting of a sequence of dialogue messages between user and assistant roles, emphasizing short sentences, high-density interactions, and particular attention to emotional expression, sense of distance, and creating a soft, intimate conversational atmosphere. Its design goal is to optimize conversational fine-tuning, suitable for supervised fine-tuning of large language models, LoRA training, and research and enhancement of Japanese dialogue models and emotional dialogue models. It is recommended for use with model series such as Gemma, Qwen, Llama, Mistral, and OpenHermes. Notes indicate that due to the datasets focus on specific conversational styles, issues like repetition, emotional overfitting, or dependency bias may arise on base models, and it is advised to mix with general dialogue data or knowledge data as needed. The dataset is released under the MIT license.





