jackf857/qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-q_t-0.45-eta-0.1-s_star-0.6-margin-log
收藏资源简介:
该数据集是从一个New-DPO训练运行中导出的每步边缘摘要统计信息。数据集包含681行训练数据,每行包含epoch、step、batch_size、mean、std、min、p10、median、p90、max、pos_frac、sample和npy等字段。数据来源于模型仓库jackf857/qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-q_t-0.45-eta-0.1-s_star-0.6,基础模型为jackf857/qwen3-8b-base-sft-hh-helpful-4xh200-batch-64-20260417-214452。训练参数包括beta为0.1,f_divergence_type为reverse_kl,f_alpha_divergence_coef为1.0,s_star为0.6,eta为0.1,q_t为0.45。数据集混合器使用了Anthropic/hh-rlhf数据集。
Per-step margin summary statistics exported from a New-DPO training run. The dataset contains 681 rows of training data, each with fields such as epoch, step, batch_size, mean, std, min, p10, median, p90, max, pos_frac, sample, and npy. The data comes from the model repository jackf857/qwen3-8b-base-new-dpo-hh-helpful-4xh200-batch-64-q_t-0.45-eta-0.1-s_star-0.6, with the base model being jackf857/qwen3-8b-base-sft-hh-helpful-4xh200-batch-64-20260417-214452. Training parameters include beta of 0.1, f_divergence_type of reverse_kl, f_alpha_divergence_coef of 1.0, s_star of 0.6, eta of 0.1, and q_t of 0.45. The dataset mixer uses the Anthropic/hh-rlhf dataset.




