Luxel/hua-qwen35-tinker-training-data
收藏资源简介:
该数据集包含为Hua风格模型有机体复制准备的Qwen3.5身份重写、Tinker渲染的训练输入,用于Non-verbal-Eval-Awareness项目。数据集包括311,083条标准化合成文档微调行和41,290条专家迭代完整追踪行,总计352,373行数据,渲染的训练令牌数为230,278,727,损失令牌数为222,091,269。数据集来源于公开的Hua等人训练数据发布,未声明新的许可证,使用时应遵守上游源条款和研究背景。
This dataset contains the Qwen3.5 identity-rewritten, Tinker-rendered training inputs prepared in the `Non-verbal-Eval-Awareness` project for a Hua-style model-organism reproduction. It includes 311,083 normalized synthetic-document fine-tuning rows and 41,290 paper-replay expert-iteration full-trace rows, totaling 352,373 rows with 230,278,727 rendered train tokens and 222,091,269 loss tokens. The dataset is derived from public Hua et al. training-data releases and does not assert a new license; use should respect the upstream source terms and research context.



