W-61/qwen3-8b-base-new-dpo-ultrafeedback-4xh200-batch-128-q_t-0.4-s_star-0.45-20260430-140517-margin
收藏资源简介:
该数据集是从一个名为New-DPO的训练运行中导出的每一步的边际摘要统计。数据集包含了训练过程中的多个特征,如epoch、step、batch_size、mean、std等统计信息,以及每个步骤的样本边际和保存的完整数组路径。训练运行基于模型W-61/qwen3-8b-base-new-dpo-ultrafeedback-4xh200-batch-128-q_t-0.4-s_star-0.45-20260430-140517,使用了HuggingFaceH4/ultrafeedback_binarized数据集进行混合训练。训练参数包括beta、f_divergence_type、f_alpha_divergence_coef、s_star、eta和q_t等。
Per-step margin summary statistics exported from a New-DPO training run. The dataset includes multiple features during the training process, such as epoch, step, batch_size, mean, std, and other statistical information, as well as per-example margins for each step and the path to the saved full array. The training run is based on the model W-61/qwen3-8b-base-new-dpo-ultrafeedback-4xh200-batch-128-q_t-0.4-s_star-0.45-20260430-140517 and uses the HuggingFaceH4/ultrafeedback_binarized dataset for mixed training. Training parameters include beta, f_divergence_type, f_alpha_divergence_coef, s_star, eta, and q_t.




