danirodriguezz/nil-ojeda-dataset-dpo
收藏官方服务:
资源简介:
这是一个用于偏好学习或对齐任务的数据集,包含423个训练示例和23个测试示例。每个示例由三个文本字段组成:prompt(提示)、rejected(被拒绝的回答)和chosen(被选择的回答),适用于训练模型区分更好和更差的响应。
This is a dataset for preference learning or alignment tasks, containing 423 training examples and 23 test examples. Each example consists of three text fields: prompt, rejected (the dispreferred response), and chosen (the preferred response), suitable for training models to distinguish between better and worse responses.
提供机构:
danirodriguezz


