literary-roleplay
收藏资源简介:
Literary Roleplay SFT是一个专门用于角色扮演模型指令微调的多语言数据集,旨在教授结构化角色扮演机制。数据集包含346个样本,覆盖英语、俄语、印地语、梵语和日语五种语言,均使用纯净词汇且方法论术语已本地化翻译。数据源自多样化的文学传统,包括果戈理、布尔加科夫、佩鲁莫夫、戈洛瓦乔夫的作品,吠陀经典(奥义书、摩诃婆罗多、罗摩衍那),经典科幻小说(阿西莫夫、克拉克、布拉德伯里、莱姆、勒吉恩、斯特鲁加茨基),经典侦探小说(柯南·道尔、克里斯蒂、哈米特、钱德勒、坡),以及日本民间传说和现代诗歌。特别包含42个女性主角样本以解决性别平衡问题。每个样本教授四种角色扮演格式之一:单轮响应、多轮场景、场景编排或情感检测。角色扮演逻辑基于三个由Super Z(GLM-5.2)编写的信标脚本:角色无关情感检测引擎v3.8、简洁与反重复引擎v3.5和场景编排引擎v2.7.18。数据集包含26个负样本(13个“坏”,13个“有缺陷”),专门展示每种反模式,用于批评者模型训练。数据样本包含14个字段,如instruction、output、format等。数据按85/15比例随机拆分为训练集(294个样本)和测试集(52个样本)。所有数据均由GLM-5.2在项目负责人编辑指导下生成,数据集采用CC0-1.0许可证发布。
Literary Roleplay SFT is a multilingual dataset specifically designed for instruction fine-tuning of role-playing models, aiming to teach structured role-playing mechanisms. The dataset contains 346 samples covering five languages: English, Russian, Hindi, Sanskrit, and Japanese, all using pure vocabulary with localized translations of methodological terms. Data is sourced from diverse literary traditions, including works by Gogol, Bulgakov, Perumov, Golovachev; Vedic classics (Upanishads, Mahabharata, Ramayana); classic science fiction (Asimov, Clarke, Bradbury, Lem, Le Guin, Strugatsky); classic detective fiction (Conan Doyle, Christie, Hammett, Chandler, Poe); and Japanese folklore and modern poetry. It specifically includes 42 female protagonist samples to address gender balance. Each sample teaches one of four role-playing formats: single-turn response, multi-turn scenario, scenario orchestration, or emotion detection. Role-playing logic is based on three beacon scripts written by Super Z (GLM-5.2): role-agnostic emotion detection engine v3.8, conciseness and anti-repetition engine v3.5, and scenario orchestration engine v2.7.18. The dataset includes 26 negative samples (13 bad, 13 defective) specifically showcasing each anti-pattern for critic model training. Data samples contain 14 fields, such as instruction, output, format, etc. The data is randomly split into training (294 samples) and test (52 samples) sets in an 85/15 ratio. All data is generated by GLM-5.2 under the editorial guidance of the project lead, and the dataset is released under the CC0-1.0 license.




