ytaek-oh/eqben-images
收藏资源简介:
--- license: apache-2.0 --- <p align="center"> <h3 align="center"><a href="https://arxiv.org/abs/2303.14465" target='_blank'> <strong>Equivariant Similarity for Vision-Language Foundation Models</strong> </a></h3> <h2 align="center">ICCV 2023</h2> <p align="center"> <a href="https://scholar.google.com/citations?hl=en&user=wFduC9EAAAAJ" target='_blank'>Tan Wang</a>, <a href="https://scholar.google.com/citations?hl=en&user=LKSy1kwAAAAJ" target='_blank'>Kevin Lin</a>, <a href="https://scholar.google.com/citations?hl=en&user=WR875gYAAAAJ" target='_blank'>Linjie Li</a>, <a href="https://scholar.google.com/citations?hl=en&user=legkbM0AAAAJ" target='_blank'>Chung-Ching Lin</a>, <a href="https://scholar.google.com/citations?hl=en&user=rP02ve8AAAAJ" target='_blank'>Zhengyuan Yang</a>, <a href="https://scholar.google.com/citations?hl=en&user=YG0DFyYAAAAJ" target='_blank'>Hanwang Zhang</a>, <a href="https://scholar.google.com/citations?hl=en&user=bkALdvsAAAAJ" target='_blank'>Zicheng Liu</a>, <a href="https://scholar.google.com/citations?hl=en&user=cDcWXuIAAAAJ" target='_blank'>Lijuan Wang</a> <br> Nanyang Technological University, Microsoft Corporation </p> </p> <br /><br /> # About This study explores the concept of equivariance in vision-language foundation models (VLMs), focusing specifically on the multimodal similarity function that is not only the major training objective but also the core delivery to support downstream tasks. Unlike the existing image-text similarity objective which only categorizes matched pairs as similar and unmatched pairs as dissimilar, equivariance also requires similarity to vary faithfully according to the semantic changes. Our key contributions are three-fold: 1. A novel benchmark named **EqBen** (Equivariant Benchmark) to benchmark VLMs with **visual-minimal change** samples. 2. A plug-and-play regularization loss **EqSim** (Equivariant Similarity Learning) to improve the equivariance of current VLMs. 3. Toolkit provides an **one-stop evaluation**: not only for EqBen, but also for previous related benchmarks (Winoground, VALSE, etc).<br> # Data Download - **Download images from huggingface hub:** Please check the Files and versions tab above. - Full-Test Set: the user can download the EqBen raw **[image data](https://drive.google.com/file/d/1e608uhd36ak_v7SnlMVaYcekBc4gBqzn/view?usp=drive_link)** (tar.gz file, ~100G) and [**annotation (after randomize)**](https://drive.google.com/file/d/1-CWEuZ5F0KQ4d94Y9rRtBsMIcqb8V7nm/view?usp=sharing) (200M) via Google Drive. **[UPDATE-2023-09]** The original annotation is the annotation after randomize (non-public) for the total fairness. And the users are required to upload the results json/np file to CodaLab for getting the final results. Due to the unstability of CodaLab, we decide to public the whole original annotation. This annotation file formalized similar to *Winoground* and can be downloaded [**here**](https://drive.google.com/file/d/1gNR4K2Cv4rbnjVRdHuBV6PuoJ5MlRXnZ/view?usp=sharing). - **Light** Full-Test Set: to improve the usability, we also provide a light version of EqBen by converting all the png image to the jpg using `convert`. Feel free to download [here](https://entuedu-my.sharepoint.com/:u:/g/personal/tan317_e_ntu_edu_sg/EcHBRcch6KREvzvGgrN67FMBUSVV4QPTQUiew0bxjcitFw?e=xiJiYL). But please note that you may make some small revisement to the path in the annotation (change the `.png` to `.jpg`). - Sub-Test Set: we also provide a 10% subset (~25K image-text pairs) for the ease of visualization and validation. The label of the EqBen sub-set is **opensource** and the **format follows the winoground style**. But please note that the samples in the subset is **randomly sorted** and not be classified to each category. Please down the raw **[image data](https://drive.google.com/file/d/13Iuirsvx34-9F_1Mjhs4Dqn59yokyUjy/view?usp=sharing)** (tar.gz file, ~10G) and [**annotation**](https://drive.google.com/file/d/18BSRf1SnBtGiEc42mzRLirXaBLzYE5Tt/view?usp=sharing) via Google Drive. --- * This is the unofficial distribution of images in the eqben benchmark. * For the official repository, please visit [https://github.com/Wangt-CN/EqBen](https://github.com/Wangt-CN/EqBen). * Some part of this README.md is taken from the official repository
license: apache-2.0 <p align="center"> <h3 align="center"><a href="https://arxiv.org/abs/2303.14465" target='_blank'> <strong>面向视觉语言基础模型的等变相似度学习</strong> </a></h3> <h2 align="center">ICCV 2023</h2> <p align="center"> <a href="https://scholar.google.com/citations?hl=en&user=wFduC9EAAAAJ" target='_blank'>王坦</a>, <a href="https://scholar.google.com/citations?hl=en&user=LKSy1kwAAAAJ" target='_blank'>林凯文</a>, <a href="https://scholar.google.com/citations?hl=en&user=WR875gYAAAAJ" target='_blank'>李林杰</a>, <a href="https://scholar.google.com/citations?hl=en&user=legkbM0AAAAJ" target='_blank'>林仲靖</a>, <a href="https://scholar.google.com/citations?hl=en&user=rP02ve8AAAAJ" target='_blank'>杨正远</a>, <a href="https://scholar.google.com/citations?hl=en&user=YG0DFyYAAAAJ" target='_blank'>张汉旺</a>, <a href="https://scholar.google.com/citations?hl=en&user=bkALdvsAAAAJ" target='_blank'>刘志成</a>, <a href="https://scholar.google.com/citations?hl=en&user=cDcWXuIAAAAJ" target='_blank'>王丽娟</a> <br> 南洋理工大学, 微软公司 </p> </p> <br /><br /> # 研究概况 本研究探讨视觉语言基础模型(Vision-Language Foundation Models, VLMs)中的等变性概念,重点关注多模态相似度函数——该函数不仅是模型的核心训练目标,亦是支撑下游任务的关键输出模块。现有图文相似度目标仅将匹配样本判定为相似、非匹配样本判定为相异,而等变性要求相似度能够忠实响应语义变化。本研究的核心贡献包含三点: 1. 提出全新基准测试集**EqBen(等变基准测试集,Equivariant Benchmark)**,通过视觉最小变化样本对视觉语言基础模型进行基准测试。 2. 提出即插即用的正则化损失函数**EqSim(等变相似度学习,Equivariant Similarity Learning)**,用于提升现有视觉语言基础模型的等变性。 3. 提供一站式评测工具包:不仅可用于EqBen的评测,还支持此前相关基准测试集(如Winoground、VALSE等)的评测。<br> # 数据下载 - **从Hugging Face Hub下载图像**:请查看页面上方的“Files and versions”标签页。 - **完整测试集**:用户可通过Google Drive下载EqBen原始**[图像数据](https://drive.google.com/file/d/1e608uhd36ak_v7SnlMVaYcekBc4gBqzn/view?usp=drive_link)**(tar.gz格式,约100GB)与**[经随机化处理的标注文件](https://drive.google.com/file/d/1-CWEuZ5F0KQ4d94Y9rRtBsMIcqb8V7nm/view?usp=sharing)**(200MB)。 **[2023-09更新]** 原始标注为经随机化处理的标注(非公开),以保证整体公平性。此前用户需将评测结果的json/np文件上传至CodaLab以获取最终评测结果。鉴于CodaLab平台存在不稳定性,我们决定公开全部原始标注文件。该标注文件的格式与*Winoground*类似,可通过[此链接](https://drive.google.com/file/d/1gNR4K2Cv4rbnjVRdHuBV6PuoJ5MlRXnZ/view?usp=sharing)下载。 - **轻量版完整测试集**:为提升数据集易用性,我们通过`convert`工具将所有PNG图像转换为JPG格式,提供EqBen的轻量版本。可通过[此链接](https://entuedu-my.sharepoint.com/:u:/g/personal/tan317_e_ntu_edu_sg/EcHBRcch6KREvzvGgrN67FMBUSVV4QPTQUiew0bxjcitFw?e=xiJiYL)自由下载。但请注意:您需要对标注文件中的图像路径进行小幅修改,将`.png`后缀更改为`.jpg`。 - **子测试集**:我们还提供了10%的子集(约25K个图文对),方便可视化与验证。EqBen子集的标注**已开源**,格式遵循Winoground风格。但请注意:子集内的样本为随机排序,未按类别划分。请通过Google Drive下载原始**[图像数据](https://drive.google.com/file/d/13Iuirsvx34-9F_1Mjhs4Dqn59yokyUjy/view?usp=sharing)**(tar.gz格式,约10GB)与**[标注文件](https://drive.google.com/file/d/18BSRf1SnBtGiEc42mzRLirXaBLzYE5Tt/view?usp=sharing)**。 --- * 本分发包为EqBen基准测试集中图像的非官方分发版本。 * 官方仓库请访问[https://github.com/Wangt-CN/EqBen](https://github.com/Wangt-CN/EqBen)。 * 本README.md的部分内容源自官方仓库。
数据集概述
关于
本研究探索了视觉-语言基础模型(VLMs)中的等变性概念,特别关注于多模态相似性函数,这不仅是主要的训练目标,也是支持下游任务的核心交付内容。与现有的仅将匹配对分类为相似、不匹配对分类为不相似的图像-文本相似性目标不同,等变性还要求相似性根据语义变化忠实地变化。我们的主要贡献有三点:
- 一个名为 EqBen(等变性基准)的新基准,用于评估具有视觉最小变化样本的VLMs。
- 一个即插即用的正则化损失 EqSim(等变性相似性学习),以提高当前VLMs的等变性。
- 工具包提供了一个一站式评估:不仅适用于EqBen,还适用于先前的相关基准(如Winoground、VALSE等)。
数据下载
- 完整测试集:用户可以通过Google Drive下载EqBen原始的**图像数据(tar.gz文件,约100G)和随机化后的标注**(200M)。
- 轻量版完整测试集:为了提高可用性,我们还提供了一个轻量版的EqBen,通过将所有png图像转换为jpg格式。可以在此处下载:轻量版数据。请注意,您可能需要对标注中的路径进行一些小的修改(将
.png改为.jpg)。 - 子测试集:我们还提供了一个10%的子集(约25K图像-文本对),以便于可视化和验证。EqBen子集的标签是开源的,格式遵循Winoground风格。请注意,子集中的样本是随机排序的,并未分类到各个类别。请在此处下载原始的**图像数据(tar.gz文件,约10G)和标注**。
- 这是eqben基准图像的非官方分发。
- 对于官方仓库,请访问https://github.com/Wangt-CN/EqBen。




