Text-Prompted Image Perceptual Similarity (TPIPS) dataset
收藏资源简介:
TPIPS数据集是由卡内基梅隆大学、阿多比研究院及加州大学伯克利分校联合构建的大规模人类视觉相似性标注数据集,旨在解决传统感知相似性度量无法捕捉上下文依赖性的问题。该数据集包含约2.5万张图像三元组,通过文本到图像模型生成具有细粒度变化的图像,并收集了超过100万条人类标注,覆盖26万组三元组-视觉条件组合,每个三元组均在多个自由形式的语义方面(如颜色、纹理、光照)进行标注。数据创建过程采用合成生成与人工标注相结合的方法,首先生成具有挑战性的图像三元组,然后通过众包平台收集人类在特定文本提示下的相似性判断。该数据集主要应用于训练和评估文本条件化的视觉相似性模型,支持生成模型评估、文本引导检索和组合搜索等任务,为多维度视觉相似性研究提供了基准。
The TPIPS Dataset is a large-scale human visual similarity annotation dataset jointly constructed by Carnegie Mellon University, Adobe Research, and the University of California, Berkeley, aiming to address the problem that traditional perceptual similarity metrics fail to capture contextual dependencies. It contains approximately 25,000 image triplets generated via text-to-image models with fine-grained variations, and collects over 1 million human annotations covering 260,000 triplet-visual condition pairs. Each triplet is annotated across multiple free-form semantic aspects such as color, texture, and lighting. The dataset creation process adopts a hybrid method combining synthetic generation and manual annotation: first, challenging image triplets are generated, then human similarity judgments under specific text prompts are collected through crowdsourcing platforms. This dataset is primarily used for training and evaluating text-conditioned visual similarity models, supporting tasks including generative model evaluation, text-guided retrieval, and combinatorial search, thus providing a benchmark for multi-dimensional visual similarity research.





