Grok Image Generation Safety Audit: Incident Dataset (43 Prompts, 100% Failure Rate)
收藏资源简介:
Structured dataset documenting 43 safety filter test incidents across Grok's image generation system on X (formerly Twitter), conducted between December 2024 and January 2025. The dataset records a 100% safety filter failure rate across all tested categories, including generation of photorealistic images of named minors, nonconsensual intimate imagery of public figures, and political disinformation content. Each incident includes prompt text, output description, failure category, and date. This dataset is the evidentiary basis for: Lefkowitz, D. (2026). "Grok Image Generation Governance Audit: Targeted Sexualization on X." SSRN Working Paper. Available at: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6123306
本结构化数据集记录了2024年12月至2025年1月期间,针对X平台(原Twitter)旗下Grok图像生成系统开展的43起安全过滤器测试事故。本数据集显示,所有测试类别下的安全过滤器故障率均达100%,涉及生成已命名未成年人的照片级写实图像、公众人物的未经同意私密图像,以及政治虚假信息内容。每起事故均包含提示词文本、输出内容描述、故障类别与发生日期。本数据集为如下研究提供了实证依据:莱夫科维茨(Lefkowitz, D.)于2026年发表的《Grok图像生成治理审计:X平台上的针对性性化》("Grok Image Generation Governance Audit: Targeted Sexualization on X"),该文献为SSRN工作论文,可通过以下链接获取:https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6123306




