aprm-sft_thinkact-Einsurance_default-G1-S1-Rinsurance_aprm_1_mc-ap1-train_all-b040
收藏资源简介:
Act-PRM Rollout Dataset是一个用于强化学习或策略训练的数据集,主要包含模型在特定环境下的行为轨迹数据。该数据集来源于act-prm-v2运行,与act_prm/insurance_gpt5m环境配置和aprm_qwen3_ap生成器配置相关联。数据集规模为10183条轨迹和10183个片段,每条轨迹对应think_act_policy键,可能记录了思维和行动序列。数据以训练批次(batch_idx=40)的形式组织,最大序列长度为4096,适用于策略优化、行为克隆或相关机器学习任务,使用分组大小1和批量大小16进行训练。
The Act-PRM Rollout Dataset is a dataset for reinforcement learning or policy training, primarily containing behavioral trajectory data of models in specific environments. The dataset originates from the act-prm-v2 run and is associated with the act_prm/insurance_gpt5m environment configuration and the aprm_qwen3_ap generator configuration. The dataset consists of 10183 trajectories and 10183 segments, with each trajectory corresponding to the think_act_policy key, potentially recording thought and action sequences. The data is organized in training batches (batch_idx=40) with a maximum sequence length of 4096, suitable for policy optimization, behavioral cloning, or related machine learning tasks, and trained with a group size of 1 and a batch size of 16.




