l0-qwen3-1.7b-compression-bs32-n16-32k-t1-no-eos-verl091-146102-rollouts
收藏资源简介:
该数据集为强化学习(RL)训练过程中生成的rollouts(轨迹响应)数据,专门用于压缩数学验证(compression MathVerify)任务。数据来自使用Qwen3-1.7B模型、批大小32、采样数16、序列长度32k、温度1、无EOS门控的训练实验。每个训练步骤提供一个经过验证的gzip压缩JSONL分片,每个分片包含512个响应。数据格式为标准JSONL,每条记录对应一个模型生成的输出。适用于强化学习中的策略优化、奖励建模或模型微调等场景。
This dataset consists of rollouts (trajectory responses) generated during reinforcement learning (RL) training, specifically for the compression MathVerify task. The data comes from training experiments using the Qwen3-1.7B model with batch size 32, sample count 16, sequence length 32k, temperature 1, and no EOS gating. Each training step provides a verified gzip-compressed JSONL shard, and each shard contains 512 responses. The data format is standard JSONL, where each record corresponds to an output generated by the model. It is suitable for policy optimization, reward modeling, or model fine-tuning in reinforcement learning scenarios.
数据集概述
基本信息
- 数据集名称:l0_Qwen3-1.7B_compression_bs32_n16_32k_t1_no_eos_verl091
- 许可证:apache-2.0
- 数据集类型:RL training rollouts(强化学习训练输出)
内容结构
- 每个训练步骤对应一个已验证的 gzip JSONL 分片。
- 每个分片包含 512 条响应。
数据处理说明
- 采用 Historical compression MathVerify,在 thinking 之后进行验证。
- 不使用 EOS gate。





