Typographic Visual Prompt Injection (TVPI) Dataset
收藏资源简介:
TVPI数据集是由香港科技大学(广州)的研究团队创建的,旨在评估不同生成模型在面对版式视觉提示注入攻击时的安全性。该数据集包含视觉语言感知(VLP)和图像到图像(I2I)两种类型的数据子集,每个子集都包括基础清晰图像、因素修改和不同目标威胁处理三个部分。VLP子集包含四个子任务,I2I子集包含两个子任务,涵盖了从类别、颜色、数量到大小等多个对象属性的识别,以及图像风格转换和全身姿态生成等任务。数据集通过在不同场景下设计特定的攻击目标,以全面探索版式视觉提示注入攻击的影响。
The TVPI Dataset was developed by the research team from The Hong Kong University of Science and Technology (Guangzhou), aiming to evaluate the safety of diverse generative models against layout visual prompt injection attacks. This dataset comprises two types of data subsets: Vision-Language Perception (VLP) and Image-to-Image (I2I). Each subset includes three components: baseline clear images, factor-modified samples, and various target threat handling modules. The VLP subset contains four subtasks, while the I2I subset has two subtasks, covering multiple object attribute recognition tasks including category, color, quantity and size, as well as tasks such as image style transfer and full-body pose generation. The dataset is designed with specific attack targets across different scenarios to comprehensively investigate the impacts of layout visual prompt injection attacks.

- 1Exploring Typographic Visual Prompts Injection Threats in Cross-Modality Generation Models香港科技大学(广州) · 2025年



