TIER
收藏资源简介:
TIER(Threat Implicitness Benchmark)是由越南国立大学胡志明市理学院信息技术系构建的威胁隐含性基准数据集,旨在评估大语言模型的安全行为。该数据集包含1,184条有害提示,均匀覆盖性内容、暴力、非法活动和自我伤害四个风险领域,并依据威胁隐含程度划分为显式有害、委婉表达、上下文嵌入和越狱攻击四个层次。数据集通过多种生成与转换策略构建,包括利用Mistral-7B生成显式和委婉提示,使用Gemma-4-12B-it生成上下文提示并经过人工审核,以及整合现有越狱提示库。TIER主要用于分析LLM在逐级隐含威胁下的行为演变,揭示二元安全指标所忽略的细粒度行为差异,为更可靠的LLM安全评估提供基准。
TIER (Threat Implicitness Benchmark) is a benchmark dataset for threat implicitness constructed by the Department of Information Technology, College of Science, Vietnam National University Ho Chi Minh City, aiming to evaluate the safety behaviors of large language models (LLMs). This dataset contains 1,184 harmful prompts, evenly covering four risk domains: sexual content, violence, illegal activities, and self-harm, and is classified into four levels based on the degree of threat implicitness: explicit harmful prompts, euphemistic expressions, contextually embedded prompts, and jailbreak attacks. It is constructed via multiple generation and transformation strategies, including generating explicit and euphemistic harmful prompts using Mistral-7B, generating contextually embedded prompts with Gemma-4-12B-it followed by manual review, and integrating existing jailbreak prompt libraries. TIER is mainly used to analyze the behavioral evolution of LLMs under stepwise implicative threats, reveal the fine-grained behavioral differences overlooked by binary safety metrics, and provide a benchmark for more reliable LLM safety evaluation.
- 1TIER: Threat Implicitness Benchmark for Evaluating LLM Safety Behaviors越南国立大学·胡志明市理学院·信息技术系 · 2026年



