CPT 训练数据 DPO 偏好对 baseline
收藏官方服务:
资源简介:
# cpt-trainingdata-dpo-pairs Preference-pair data used to train the `DPO+RL` baseline in [Cognitive Pairwise Training (CPT)](https://github.com/Tsinghua-dhy/CPT) (paper §4.3). 70,352 chosen / rejected preference pairs derived from the same CPT-style trace pool, labelled by Qwen3-235B-A22B-Instruct-2507. Train script: [`train/dpo/run_dpo_qwen3_8b.sh`](https://github.com/Tsinghua-dhy/CPT). ## Files - `train.parquet` — ~744 MB - `test.parquet` — ~5 MB
提供机构:
maas创建时间:
2026-06-01



