FuseLIP
收藏资源简介:
FuseLIP是一个新型的多模态嵌入模型,基于早期融合的离散图像和文本标记,通过单个Transformer编码器进行处理,实现了超越或可比于现有晚期融合方法的性能。该模型可以有效地在单模态和多模态数据上进行训练,使用标准的对比损失并结合硬负样本,同时支持并显著受益于掩码建模目标。FuseLIP旨在解决多模态输入的编码问题,即如何将图像和文本对编码为单个特征向量,从而在视觉语言对齐和零样本任务中取得优异表现。
FuseLIP is a novel multimodal embedding model that processes early-fused discrete image and text tokens via a single Transformer encoder, achieving performance that exceeds or matches that of existing late-fusion methods. This model can be efficiently trained on both unimodal and multimodal data, using standard contrastive loss combined with hard negatives, while supporting and significantly benefiting from masked modeling objectives. FuseLIP aims to solve the encoding problem of multimodal inputs: how to encode image-text pairs into a single feature vector, thereby achieving excellent performance in vision-language alignment and zero-shot tasks.

- 1FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens图宾根大学图宾根人工智能中心 · 2025年



