ClearCLIP: Decomposing CLIP Representations for Dense Vision-Language Inference

Name: ClearCLIP: Decomposing CLIP Representations for Dense Vision-Language Inference
Creator: DR-NTU (Data)
Published: 2025-10-10 04:59:28
License: 暂无描述

DataCite Commons2025-10-10 更新2025-04-16 收录

下载链接：

https://researchdata.ntu.edu.sg/citation?persistentId=doi:10.21979/N9/S6NTDJ

下载链接

链接失效反馈

官方服务：

资源简介：

Despite the success of large-scale pretrained Vision-Language Models (VLMs) especially CLIP in various open-vocabulary tasks, their application to semantic segmentation remains challenging, producing noisy segmentation maps with mis-segmented regions. In this paper, we carefully re-investigate the architecture of CLIP, and identify residual connections as the primary source of noise that degrades segmentation quality. With a comparative analysis of statistical properties in the residual connection and the attention output across different pretrained models, we discover that CLIP’s image-text contrastive training paradigm emphasizes global features at the expense of local discriminability, leading to noisy segmentation results. In response, we propose ClearCLIP, a novel approach that decomposes CLIP’s representations to enhance open-vocabulary semantic segmentation. We introduce three simple modifications to the final layer: removing the residual connection, implementing the self-self attention, and discarding the feed-forward network. ClearCLIP consistently generates clearer and more accurate segmentation maps and outperforms existing approaches across multiple benchmarks, affirming the significance of our discoveries.

提供机构：

DR-NTU (Data)

创建时间：

2024-09-25

5,000+

优质数据集

54 个

任务类型

进入经典数据集