STAR-1
收藏资源简介:
STAR-1是一个针对大型推理模型(LRMs)的安全数据集,由加州大学圣塔克鲁兹分校、谷歌和劳伦斯利弗莫尔国家实验室共同创建。该数据集包含1000个样本,旨在满足LRMs在安全对齐方面的关键需求。STAR-1的构建始于一个包含41K个安全训练样本的多样化数据集,然后利用深思推理范式结构化数据,并最终通过评分筛选降低至1K个样本。该数据集在保障安全性的同时,也注重保持样本的多样性,适用于提升LRMs在安全关键场景下的推理稳健性和可靠性。
STAR-1 is a safety-oriented dataset tailored for Large Reasoning Models (LRMs), jointly developed by the University of California, Santa Cruz, Google, and the Lawrence Livermore National Laboratory. This dataset contains 1,000 samples, designed to address the critical needs of LRMs regarding safety alignment. The construction of STAR-1 starts with a diverse dataset comprising 41,000 safety training samples, structures the data using the deliberative reasoning paradigm, and finally narrows it down to 1,000 samples through scoring-based filtering. While ensuring safety, this dataset also prioritizes maintaining sample diversity, and is applicable to improving the reasoning robustness and reliability of LRMs in safety-critical scenarios.




