mkd-chanwoo/keural-conversation-ko
收藏资源简介:
keural-conversation是一个韩语日常对话AI学习用的合成数据集,基于9个主题(包括自我介绍及问候、情感表达与共情、天气、季节、日常近况、兴趣、爱好、喜欢的事物、创意写作、意见与讨论、幽默与笑话、翻译与校对、职场礼仪)和1,350个场景,生成了191,093个多轮对话对。对话结构为2轮(用户→助手→用户→助手),使用google/gemma-4-26B-A4B-it模型生成。原始生成数为202,435个,接受率为99.97%,重复去除率为5.60%。数据集以JSON格式存储,包含ID、主题、场景描述、任务类型、领域、模型名称、轮次和序列化对话等字段。数据通过四阶段API调用生成,并经过质量过滤和重复去除处理。该数据集旨在用于AI模型训练,是合成数据而非真实人类对话。
keural-conversation is a synthetic dataset for AI learning of Korean daily conversations, based on 9 topics (including self-introduction and greetings, emotional expression and empathy, weather, season, daily updates, hobbies, interests, likes, creative writing, opinions and discussions, humor and jokes, translation and proofreading, workplace etiquette) and 1,350 scenarios, generating 191,093 multi-turn dialogue pairs. The conversation structure is 2-turn (user → assistant → user → assistant), generated using the google/gemma-4-26B-A4B-it model. The original generation count is 202,435, with an acceptance rate of 99.97% and a deduplication rate of 5.60%. The dataset is stored in JSON format, including fields such as ID, topic, situation description, task type, domain, model name, turns, and serialized conversation. Data is generated through a four-stage API call process and undergoes quality filtering and deduplication. This dataset is intended for AI model training and is synthetic data, not real human conversations.



