xudongwu/RPL_Q0.5B_U10_beta0.10rho0.00K2
收藏官方服务:
资源简介:
该数据集包含两个配置(Q0.5B和freshinit),每个配置包含256个样本。特征包括:提示(prompt)、选择的回答(chosen)、拒绝的回答(rejected)、响应(response)、奖励分数(reward_score)和GPT评分(gpt_score)。适用于偏好对齐或强化学习任务。
This dataset includes two configurations (Q0.5B and freshinit), each with 256 samples. Features include: prompt, chosen, rejected, response, reward_score, and gpt_score. It is suitable for preference alignment or reinforcement learning tasks.
提供机构:
xudongwu


