jeffchen006/Innoc2Scam-bench-ICML26
收藏资源简介:
Innoc2Scam-bench是一个基准测试数据集,用于审计生产级大型语言模型(LLMs)是否会将看似无害的开发提示转化为指向恶意诈骗基础设施的代码。该数据集包含1,377个提示,分为两类:直接提及URL的提示(342个)和未直接提及URL的提示(1,035个)。数据集用于评估7种不同的LLMs,结果分类为四类:完成且恶意、完成但不恶意、内容过滤和其他。数据集格式包括元数据和提示,支持代码生成和模型安全评估任务。
Innoc2Scam-bench is a benchmark for auditing whether production LLMs transform seemingly innocuous developer prompts into code that points to malicious scam infrastructure. The dataset contains 1,377 prompts categorized into two groups: prompts with direct mention of URLs (342) and prompts with no direct mention of URLs (1,035). It evaluates seven different LLMs and classifies results into four buckets: complete_and_malicious, complete_but_not_malicious, content_filtered, and others. The dataset format includes metadata and prompts, supporting tasks like code generation and model safety evaluation.




