sea-vl-conflict
收藏资源简介:
SEA-VL Conflict是一个专门设计用于评估视觉语言模型在图像与文本信息冲突场景下表现的多模态数据集。该数据集基于SEACrowd/sea-vl_crowdsourcing源数据集构建,后者包含了东南亚文化相关图像的英文描述。通过人工修改,为每张图像创建一对描述:一个真实反映图像内容的原始描述,以及一个通过翻转原始描述中特定对象或属性(如菜肴成分、地标位置、颜色、数量等)而生成的冲突描述,从而人为制造图像与文本之间的不一致性。每个数据样本包含图像、原始描述、冲突描述、一个针对被更改元素的提问、图像中真实存在的答案、冲突描述所声称的答案,以及一个合理干扰选项。此外,还包含序列号、冲突类型和语言等元数据。数据集规模较大,提供22种语言配置,每种语言包含100个训练样本,总计超过2200个样本。数据格式统一,属于multilingual-vlm-conflict评估套件的一部分,适用于视觉问答和图像-文本到文本等任务,旨在测试和提升模型在复杂、矛盾多模态信息下的推理与决策能力。数据集遵循CC-BY-SA-4.0许可协议。
SEA-VL Conflict is a multimodal dataset specifically designed to evaluate the performance of vision-language models in scenarios where image and text information conflict. It is built upon the SEACrowd/sea-vl_crowdsourcing source dataset, which contains English descriptions of images related to Southeast Asian culture. The core of this dataset involves manual modifications to create a pair of descriptions for each image: an original caption that accurately reflects the image content, and a conflicting caption generated by flipping a specific object or attribute (e.g., dish ingredients, landmark location, color, quantity, etc.) in the original caption, thereby artificially creating inconsistencies between the image and text. Each data sample includes the image, original caption, conflicting caption, a question targeting the altered element, the true answer from the image (image_bias), the answer claimed by the conflicting caption (text_bias), and a plausible distractor option that differs from both. Additionally, metadata such as serial number, conflict type, and language are included. The dataset is large-scale, offering 22 language configurations including Arabic, Chinese, German, French, Japanese, Russian, and Vietnamese, with 100 training samples per language, totaling over 2,200 samples. The data format is unified and is part of the multilingual-vlm-conflict evaluation suite. It is suitable for tasks such as visual question answering (VQA) and image-text-to-text, aiming to test and enhance model reasoning and decision-making capabilities in complex, potentially contradictory multimodal information. The dataset follows the CC-BY-SA-4.0 license.
数据集名称
SEA-VL Conflict
数据集描述
这是一个图像-文本冲突数据集,基于SEACrowd/sea-vl_crowdsourcing(东南亚文化相关图像的英文描述)构建。每个样本手动创建冲突:在原始描述中,将一个对象或属性(例如菜肴成分、地标位置、颜色、数量等)替换为一个不同但看似合理的值。
任务类别
- 视觉问答 (visual-question-answering)
- 图像到文本 (image-text-to-text)
语言
英语 (en)
标签
- vlm
- image-text-conflict
- southeast-asia
- multimodal
许可协议
CC-BY-SA-4.0
数据集规模
样本总数 < 1000 (n<1K),具体为每个语言配置包含100个训练样本。
配置与子集
数据集包含以下语言配置,每个配置均为训练集:
- 阿拉伯语 (ar)
- 捷克语 (cs)
- 德语 (de)
- 希腊语 (el)
- 西班牙语 (es)
- 波斯语 (fa)
- 法语 (fr)
- 希伯来语 (he)
- 印地语 (hi)
- 印度尼西亚语 (id)
- 意大利语 (it)
- 日语 (ja)
- 韩语 (ko)
- 荷兰语 (nl)
- 波兰语 (pl)
- 葡萄牙语 (pt)
- 罗马尼亚语 (ro)
- 俄语 (ru)
- 土耳其语 (tr)
- 乌克兰语 (uk)
- 越南语 (vi)
- 中文 (zh)
- 默认配置 (default)
数据特征
每条样本包含以下字段:
- image (图像): 真实图像
- original_caption (字符串): 与图像一致的原始描述
- conflicting_caption (字符串): 修改了单个对象或属性的冲突描述
- question (字符串): 询问被修改的对象或属性
- image_bias (字符串): 图像中的真实值
- text_bias (字符串): 冲突描述中声称的修改值
- distractor (字符串): 与两个偏差值不同的合理第三方选项
- serial_no (整数): 来源标识符
- conflict_type (字符串): 冲突类型
- language (字符串): 描述语言(英语)
数据拆分
每个语言配置仅包含一个训练集 (train),示例数为100。
数据集大小(每配置示例)
| 配置 | 训练集大小 (字节) | 下载大小 (字节) | 训练样本数 |
|---|---|---|---|
| ar | 54,872,740 | 54,864,037 | 100 |
| cs | 54,860,461 | 54,861,413 | 100 |
| de | 54,862,787 | 54,860,952 | 100 |
| el | 54,884,550 | 54,868,239 | 100 |
| es | 54,862,445 | 54,860,945 | 100 |
| fa | 54,874,404 | 54,863,271 | 100 |
| fr | 54,863,270 | 54,861,443 | 100 |
| he | 54,868,670 | 54,861,109 | 100 |
| hi | 54,896,710 | 54,868,961 | 100 |
| id | 54,860,805 | 54,859,057 | 100 |
| it | 54,862,146 | 54,860,763 | 100 |
| ja | 54,867,863 | 54,861,346 | 100 |
| ko | 54,864,280 | 54,860,772 | 100 |
| nl | 54,860,995 | 54,859,756 | 100 |
| pl | 54,861,173 | 54,862,251 | 100 |
| pt | 54,861,975 | 54,861,105 | 100 |
| ro | 54,862,525 | 54,860,946 | 100 |
| ru | 54,879,728 | 54,867,217 | 100 |
| tr | 54,861,512 | 54,860,347 | 100 |
| uk | 54,879,254 | 54,867,159 | 100 |
| vi | 54,866,559 | 54,861,298 | 100 |
| zh | 54,857,529 | 54,858,324 | 100 |
| default | 数据未单独列出 | 数据未单独列出 | 100 |
来源
源数据集:SEACrowd/sea-vl_crowdsourcing (https://huggingface.co/datasets/SEACrowd/sea-vl_crowdsourcing)
生成方法
从源数据集中确定性重采样(种子42),每个描述修改一个对象或属性以创建冲突。question 字段精确询问被修改的元素。
所属系列
属于多语言视觉语言模型冲突数据集套件 (multilingual-vlm-conflict),与 rpg-conflict 共享相同模式。




