相关数据集
OpenRLHF/preference_dataset_mixture2_and_safe_pku
--- {} --- > Copy from https://huggingface.co/datasets/weqweasdas/preference_dataset_mixture2_and_safe_pku # Reward Model Overview This
Hugging Face2024-06-14 更新1350
OpenRLHF/preference_700K
该数据集是一个增强版本,包含两个主要部分:rejected和chosen,每个部分都包含content和role两个字段,数据类型为字符串。此外,数据集还包含rejected_score和chosen_score两个字段,数据类型为float64。数据集分为一个训练集,包含700,000个样本,总大小为2,802,733,004字节。
Hugging Face2024-07-13 更新190
OpenRLHF/dapo-math-17k
--- dataset_info: features: - name: prompt list: - name: content dtype: string - name: role dtype: string - name: label dtype: string splits: - name: train nu
Hugging Face2026-02-06 更新100
OpenRLHF/prompt-collection-v0.1-dev-100k
这个5k的数据集是从https://huggingface.co/datasets/RLHFlow/prompt-collection-v0.1中抽取的,保持了与完整数据集相似的分布,并且仅用于OpenRLHF中的PPO训练开发和验证。同时,尊重并保留了原始数据提供者的所有权利。
Hugging Face2024-12-13 更新470
OpenRLHF/aime-2024
--- dataset_info: features: - name: prompt list: - name: content dtype: string - name: role dtype: string - name: label dtype: string splits: - name: train nu
Hugging Face2026-02-06 更新130



