DGM4+
收藏资源简介:
DGM4+数据集是为了解决现有数据集在处理全局场景不一致性方面的不足而创建的。它扩展了DGM4数据集,增加了5000个高质量样本,引入了前景-背景(FG-BG)不匹配及其与文本操纵的混合样本。数据集使用OpenAI的gpt-image模型生成人本主义新闻风格图像,其中真实的人物被放置在荒谬或不可能的背景下。数据集提供了三种形式的字幕:字面意思、文本属性和文本分割,产生了三种新的操纵类别:FG-BG、FG-BG+TA和FG-BG+TS。数据集还进行了严格的质量控制,包括可见面孔、感知哈希去重、OCR文本清除和现实标题长度。DGM4+数据集旨在加强多模态模型(如HAMMER)的评价,这些模型目前难以处理FG-BG不一致性。
The DGM4+ dataset was developed to address the shortcomings of existing datasets in handling global scene inconsistencies. It extends the original DGM4 dataset by adding 5,000 high-quality samples, and introduces foreground-background (FG-BG) mismatches and hybrid samples incorporating text manipulation. The dataset uses OpenAI’s GPT-image model to generate humanistic news-style images, where real human subjects are placed in absurd or impossible contexts. It offers three types of captions: literal meaning, text attributes, and text segmentation, leading to three new manipulation categories: FG-BG, FG-BG+TA, and FG-BG+TS. Strict quality control is conducted, including verification of visible human faces, perceptual hash deduplication, OCR text cleanup, and compliance with realistic caption lengths. The DGM4+ dataset aims to strengthen the evaluation of multimodal models such as HAMMER, which currently face challenges in handling FG-BG inconsistency.

- 1DGM4+: Dataset Extension for Global Scene Inconsistency麻省理工学院计算机科学与人工智能实验室 · 2025年



