jackf857/qwen3-8b-base-new-dpo-hh-harmless-4xh200-batch-64-q_t-0.45-eta-0.1-s_star-0.8-margin-log
收藏资源简介:
该数据集是从一个名为New-DPO的训练运行中导出的每一步的边缘摘要统计信息。数据集包含了训练过程中的多个统计特征,如epoch、step、batch_size、mean、std、min、p10、median、p90、max、pos_frac、sample和npy等。数据集来源于模型仓库jackf857/qwen3-8b-base-new-dpo-hh-harmless-4xh200-batch-64-q_t-0.45-eta-0.1-s_star-0.8,基础模型为jackf857/qwen3-8b-base-sft-hh-harmless-4xh200-batch-64-20260417-214452。训练运行名称为qwen3-8b-base-new-dpo-hh-harmless-4xh200-batch-64-q_t-0.45-eta-0.1-s_star-0.8,W&B项目为qwen3-hh-new-dpo-hyperparamter-sweep。边缘训练参数包括beta为0.1,f_divergence_type为reverse_kl,f_alpha_divergence_coef为1.0,s_star为0.8,eta为0.1,q_t为0.45。数据集混合器使用了Anthropic/hh-rlhf数据集,比例为1.0。
This dataset contains per-step margin summary statistics exported from a New-DPO training run. The dataset includes various statistical features during training, such as epoch, step, batch_size, mean, std, min, p10, median, p90, max, pos_frac, sample, and npy. The dataset originates from the model repository jackf857/qwen3-8b-base-new-dpo-hh-harmless-4xh200-batch-64-q_t-0.45-eta-0.1-s_star-0.8, with the base model being jackf857/qwen3-8b-base-sft-hh-harmless-4xh200-batch-64-20260417-214452. The training run name is qwen3-8b-base-new-dpo-hh-harmless-4xh200-batch-64-q_t-0.45-eta-0.1-s_star-0.8, and the W&B project is qwen3-hh-new-dpo-hyperparamter-sweep. Margin training arguments include beta of 0.1, f_divergence_type of reverse_kl, f_alpha_divergence_coef of 1.0, s_star of 0.8, eta of 0.1, and q_t of 0.45. The dataset mixer uses the Anthropic/hh-rlhf dataset with a ratio of 1.0.




