swiss-ai/MaxMin_Tr_3600-Filtered-Decontaminated
收藏资源简介:
Apertus 1.5偏好数据集是用于Apertus v1.5对齐训练(在线DPO)的最终偏好数据集。该数据集源自Ai2的Olmo 3 Dolci-Instruct-DPO数据集,但仅使用了其中的提示(prompts);所有“chosen”和“rejected”响应均由数据集构建者生成。构建过程包括:提示取自Dolci-Instruct-DPO,响应通过swiss-ai/posttraining-data仓库和指定模型池生成,应用了去污染过滤(针对评估套件)和长度过滤(去除提示与生成完成超过3600个令牌的行)。经过过滤后,剩余245,898个提示。数据集基于ODC-BY许可证发布,适用于研究和教育用途,遵循Ai2的负责任使用指南。
This is the final preference dataset used for Apertus v1.5 alignment training (online DPO). It is derived from Ai2s Olmo 3 Dolci-Instruct-DPO dataset. In the online DPO setting, only the prompts are reused from Dolci-Instruct-DPO; all chosen and rejected responses in this dataset were generated by the creators. The dataset was built by taking prompts from Dolci-Instruct-DPO, generating responses using the swiss-ai/posttraining-data repository and a specified model pool, applying decontamination filtering against an evaluation suite, and removing rows where the prompt plus generated completions exceeded 3600 tokens. After decontamination and length filtering, 245,898 prompts remain. It is released under the ODC-BY license for research and educational use in accordance with Ai2s Responsible Use Guidelines.




