W-61/qwen3-8b-base-new-dpo-ultrafeedback-4xh200-batch-128-q_t-0.43-s_star-0.3-20260430-192039-margin
收藏资源简介:
该数据集是从一个New-DPO训练运行中导出的每步边际摘要统计。它包含了训练过程中的多个统计特征,如epoch、step、batch_size、mean、std等,以及每步的边际样本和可选的完整边际数组保存路径。数据集的分割为train,包含477个例子。源运行信息包括模型仓库ID、基础模型、训练运行名称等。边际训练参数包括beta、f_divergence_type、s_star等。数据集混合器使用了HuggingFaceH4/ultrafeedback_binarized。
Per-step margin summary statistics exported from a New-DPO training run. It includes various statistical features during training such as epoch, step, batch_size, mean, std, etc., as well as per-step margin samples and optional paths to saved full margin arrays. The dataset split is train, containing 477 examples. Source run information includes model repo id, base model, training run name, etc. Margin training arguments include beta, f_divergence_type, s_star, etc. The dataset mixer uses HuggingFaceH4/ultrafeedback_binarized.




