AIMS (Annotated Intents for Model Safety)
收藏资源简介:
AIMS(模型安全标注意图)数据集是由格罗宁根大学等机构构建的面向大语言模型安全分类的人类标注意图数据集,旨在通过显式建模用户意图来提升安全分类器的鲁棒性。该数据集包含1,724条经过筛选的困难安全提示,每条数据均配有单句意图描述和伤害标签,数据源自WildGuardMix数据集中通过集成模型不确定性估计选出的模糊、对抗性及边界案例。其构建过程涉及基于概率预测的候选提示筛选,并由标注者人工推断用户潜在意图并标注伤害等级,最终经质量过滤形成1,275条唯一提示对。该数据集主要应用于大语言模型安全防护领域,通过提供意图中心化的监督信号,支持监督微调、偏好学习、推理蒸馏和强化学习等多种训练机制,以解决传统安全分类器因忽略深层用户目标而导致的误判或过度拒绝问题。
AIMS (Model Safety Annotation Intent) dataset is a human-annotated intent dataset for large language model (LLM) safety classification, constructed by the University of Groningen and other institutions. It aims to improve the robustness of safety classifiers by explicitly modeling user intent. This dataset contains 1,724 filtered challenging safety prompts, each paired with a single-sentence intent description and harm label. The data is sourced from ambiguous, adversarial, and borderline cases selected via ensemble model uncertainty estimation from the WildGuardMix dataset. Its construction process involves candidate prompt screening based on probabilistic predictions, followed by human annotators inferring users' latent intents and annotating harm levels, ultimately resulting in 1,275 unique prompt pairs after quality filtering. This dataset is primarily applied in the field of LLM safety protection, providing intent-centric supervision signals to support various training mechanisms such as supervised fine-tuning, preference learning, inference distillation, and reinforcement learning, so as to address the misjudgment or over-rejection issues of traditional safety classifiers caused by their neglect of deep-seated user goals.





