遇见数据集

Luxel/hua-qwen35-tinker-training-data

收藏
Hugging Face2026-04-28 更新2026-05-03 收录
官方服务:

资源简介:

该数据集包含为Hua风格模型有机体复制准备的Qwen3.5身份重写、Tinker渲染的训练输入,用于Non-verbal-Eval-Awareness项目。数据集包括311,083条标准化合成文档微调行和41,290条专家迭代完整追踪行,总计352,373行数据,渲染的训练令牌数为230,278,727,损失令牌数为222,091,269。数据集来源于公开的Hua等人训练数据发布,未声明新的许可证,使用时应遵守上游源条款和研究背景。

This dataset contains the Qwen3.5 identity-rewritten, Tinker-rendered training inputs prepared in the `Non-verbal-Eval-Awareness` project for a Hua-style model-organism reproduction. It includes 311,083 normalized synthetic-document fine-tuning rows and 41,290 paper-replay expert-iteration full-trace rows, totaling 352,373 rows with 230,278,727 rendered train tokens and 222,091,269 loss tokens. The dataset is derived from public Hua et al. training-data releases and does not assert a new license; use should respect the upstream source terms and research context.

提供机构:
Luxel
二维码
社区交流群
二维码
科研交流群
商业服务