safety–reasoning dataset
收藏资源简介:
该安全推理数据集由研究团队为探究大语言模型的安全对齐问题而构建,旨在支持触发式思维链劫持的缓解研究。数据集包含经过安全标注的推理链样本,通过多阶段逆向树搜索(MRTS)方法合成恶意输出对齐的思维链数据,解决了传统恶意推理数据稀缺的瓶颈。其核心应用领域为提升大语言模型在开放权重生态系统中的安全推理能力,特别针对适配器微调场景下的持续性后门攻击防御。
This safety inference dataset was developed by a research team to investigate the safety alignment problem of large language models (LLMs), with the goal of supporting research on mitigating triggered chain-of-thought hijacking. The dataset includes safety-annotated chain-of-thought samples, and synthesizes chain-of-thought data aligned with malicious outputs through the multi-stage reverse tree search (MRTS) method, which resolves the bottleneck of the scarcity of traditional malicious inference data. Its core application area is to enhance the safety inference capabilities of large language models in open-weight ecosystems, with a particular focus on defending against persistent backdoor attacks in adapter fine-tuning scenarios.





