leduo123/assignment4-preferences
收藏官方服务:
资源简介:
该数据集是为作业4创建的。它包含了50个偏好训练示例,这些示例是从LIMA数据集中采样的指令构建的。对于每个指令,使用`Qwen/Qwen2.5-7B-Instruct`生成了5个候选响应,并使用`llm-blender/PairRM`对这些响应进行了排名。排名最高的响应被用作`chosen`,排名最低的响应被用作`rejected`。
This dataset was created for Assignment 4. It contains 50 preference training examples built from instructions sampled from the LIMA dataset. For each instruction, 5 candidate responses were generated using `Qwen/Qwen2.5-7B-Instruct`, and the responses were ranked with `llm-blender/PairRM`. The highest-ranked response was used as `chosen`, and the lowest-ranked response was used as `rejected`.
提供机构:
leduo123


