遇见数据集

CPT 训练数据 DPO 偏好对 baseline

收藏
魔搭社区2026-06-10 更新2026-07-15 收录
官方服务:

资源简介:

# cpt-trainingdata-dpo-pairs Preference-pair data used to train the `DPO+RL` baseline in [Cognitive Pairwise Training (CPT)](https://github.com/Tsinghua-dhy/CPT) (paper §4.3). 70,352 chosen / rejected preference pairs derived from the same CPT-style trace pool, labelled by Qwen3-235B-A22B-Instruct-2507. Train script: [`train/dpo/run_dpo_qwen3_8b.sh`](https://github.com/Tsinghua-dhy/CPT). ## Files - `train.parquet` — ~744 MB - `test.parquet` — ~5 MB

提供机构:
maas
创建时间:
2026-06-01
二维码
社区交流群
二维码
科研交流群
商业服务