Thought-Aligner训练数据集
收藏资源简介:
Thought-Aligner训练数据集是用于训练Thought-Aligner模型的数据集。该数据集包含5000条指令,涵盖了十个典型的场景,能够广泛代表智能体的能力和工具集。数据集通过模拟ReAct执行轨迹生成,包含超过11400个安全和不安全的思维对。数据集的构建过程结合了LLM辅助生成和人工验证,以确保质量和准确性。数据集用于微调Thought-Aligner-1.5B和Thought-Aligner-7B模型,并在三个智能体安全基准上部署。实验结果表明,这两个模型将智能体的行为安全提高到平均90%,显示出显著的安全性能提升。
The Thought-Aligner training dataset is designed for training the Thought-Aligner model. It contains 5000 instructions covering ten typical scenarios, which broadly represent the capabilities and toolkits of AI Agents. Generated by simulating ReAct execution trajectories, the dataset includes over 11,400 pairs of safe and unsafe thought processes. The dataset's construction combines LLM-assisted generation and manual verification to ensure its quality and accuracy. This dataset is used for fine-tuning the Thought-Aligner-1.5B and Thought-Aligner-7B models, and has been deployed on three agent safety benchmarks. Experimental results show that these two models improve the behavioral safety of agents to an average of 90%, demonstrating significant improvements in safety performance.

- 1Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought Correction复旦大学 上海创新研究院 · 2025年



