xudongwu/RPL_Q7B_U10_beta0.10rho0.05K4_sf1.00
收藏资源简介:
该数据集名为Q7B,包含256个示例,用于训练或评估对话系统或偏好对齐模型。每个示例包括一个提示(prompt)、选择的响应(chosen)、拒绝的响应(rejected)、响应(response)以及两个评分字段:奖励分数(reward_score,浮点类型)和GPT分数(gpt_score,浮点类型)。这些特征可能用于比较不同响应的质量,支持强化学习或模型优化任务。数据集大小为约1.49 MB,下载大小为约0.79 MB,仅包含一个默认分割。
This dataset, named Q7B, contains 256 examples for training or evaluating dialogue systems or preference alignment models. Each example includes a prompt, a chosen response, a rejected response, a response, and two scoring fields: reward_score (float64) and gpt_score (float64). These features are likely used to compare the quality of different responses, supporting tasks such as reinforcement learning or model optimization. The dataset size is approximately 1.49 MB, with a download size of about 0.79 MB, and includes only a default split.




