JollyFraud/crimeopus-distill-v2
收藏资源简介:
CrimeOpus 4.7蒸馏编码数据集是一个用于微调CrimeOpus 4.7-v2(LoRA训练)的数据集。它包含来自多个来源的427条数据,包括DeepSeek-Chat蒸馏(157条,多领域编码/推理)、Git commit-diff(91条,真实代码库模式)和未审查种子(179条,无拒绝的有用性)。数据格式为ChatML消息数组,包含系统、用户和助手的角色内容。数据集统计信息显示,平均用户消息长度约为281字符,平均助手响应长度约为2745字符,总令牌数估计为323k。数据集支持意大利语和英语,以及代码内容,用于QLoRA微调。
CrimeOpus 4.7 — Distilled Coding Dataset is a fine-tuning dataset for CrimeOpus 4.7-v2 (LoRA training). It contains 427 entries from multiple sources, including DeepSeek-Chat distillation (157, multi-domain coding/reasoning), Git commit-diff (91, real codebase patterns), and uncensored seed (179, refusal-free helpfulness). The data format is a ChatML messages array with system, user, and assistant roles. Dataset statistics show an average user message length of ~281 chars, average assistant response length of ~2745 chars, and total estimated tokens of ~323k. The dataset supports Italian and English languages, as well as code content, and is used for QLoRA fine-tuning.




