RLAIF/dpo_thinking_reddit_judge_1e-6_0.02_4B_4B_with_gold_labels_kl_estimation
收藏数据链接:
官方服务:
资源简介:
--- dataset_info: features: - name: step dtype: int64 - name: question dtype: string - name: ref dtype: string - name: current dtype: string - name: ref_current dtype: int64 - name: current_ref dtype: int64 - name: kl_divergence dtype: float64 - name: gold dtype: int64 - name: kl_estimated dtype: float64 splits: - name: train num_bytes: 237908256 num_examples: 27000 download_size: 127422141 dataset_size: 237908256 configs: - config_name: default data_files: - split: train path: data/train-* ---
提供机构:
RLAIF


