W-61/llama-3-8b-base-new-dpo-hh-harmless-4xh200-batch-64-q_t-0.45-s_star-0.4-eta-0.5-margin-log
收藏资源简介:
该数据集是从一个New-DPO训练运行中导出的每一步的边缘统计摘要。包含训练过程中的多个统计特征,如epoch(训练轮次)、step(步骤)、batch_size(批次大小)、mean(平均值)、std(标准差)、min(最小值)、p10(第10百分位数)、median(中位数)、p90(第90百分位数)、max(最大值)、pos_frac(正分数)、sample(每个步骤的有效批次的边缘值样本)和npy(可选,保存完整边缘数组的路径)。数据集来源于特定的模型训练运行,使用了Anthropic/hh-rlhf数据集进行混合训练。
This dataset is a per-step marginal statistical summary derived from a New-DPO training run. It includes multiple statistical features collected during the training process, such as epoch (training epoch), step (training step), batch_size (batch size), mean (mean value), std (standard deviation), min (minimum value), p10 (10th percentile), median (median value), p90 (90th percentile), max (maximum value), pos_frac (positive fraction), sample (samples of marginal values from valid batches per step), and npy (optional, path to save the complete marginal array). This dataset originates from a specific model training run that utilized the Anthropic/hh-rlhf dataset for mixed training.



