DiffCap-Bench
收藏资源简介:
DiffCap-Bench是由华南理工大学、腾讯混元等机构联合构建的图像差异描述基准数据集,包含1075组高质量图像对,覆盖物体、属性、动作等十类差异维度,总计6713条人工验证的原子级差异项。数据通过多源采样(网页图像、广告等)与合成生成(2D/3D技术)相结合构建,并经过严格的质量过滤。该数据集旨在评估多模态大模型在细粒度图像差异描述任务中的性能,尤其关注语义一致性与幻觉抑制能力,为图像编辑流水线提供可靠的差异分析基准。
DiffCap-Bench is a benchmark dataset for image difference description jointly constructed by South China University of Technology, Tencent Hunyuan and other institutions. It contains 1075 high-quality image pairs, covering ten difference dimensions such as objects, attributes and actions, with a total of 6713 manually verified atomic-level difference items. The dataset is built through the combination of multi-source sampling (web images, advertisements, etc.) and synthetic generation (2D/3D technologies), and has undergone strict quality filtering. This dataset aims to evaluate the performance of multimodal large language models in the fine-grained image difference description task, with particular focus on semantic consistency and hallucination suppression capabilities, providing a reliable difference analysis benchmark for image editing pipelines.
根据提供的数据集详情页面地址和README文件内容,该数据集详情页目前仅显示“comming soon”,尚无任何与数据集相关的具体信息,因此无法进行有效的总结和概述。




