VQ-VA World
收藏资源简介:
VQ-VA World是由字节跳动种子等机构联合构建的大规模视觉问答数据集,专注于提升模型的视觉问题-视觉回答能力。该数据集包含180万高质量交错图像-文本样本,涵盖世界知识、设计知识和推理三大领域,通过智能代理流水线从网络文档中挖掘具有语义关联的图像对并生成自然语言问题。该数据集旨在解决开放域视觉生成任务中的知识推理瓶颈,推动开源模型在需要现实世界知识和多步推理的视觉问答任务上的发展。
VQ-VA World is a large-scale visual question answering (VQA) dataset jointly constructed by ByteDance Seed and other institutions, which focuses on enhancing the visual question-to-visual answer capabilities of models. This dataset contains 1.8 million high-quality interleaved image-text samples, spanning three core domains: world knowledge, design knowledge, and reasoning. It extracts semantically related image pairs from web documents through an intelligent agent pipeline and generates corresponding natural language questions. The dataset aims to address the knowledge reasoning bottlenecks in open-domain visual generation tasks, and advance the development of open-source models for visual question answering tasks that require real-world knowledge and multi-step reasoning.
VQ-VA World 数据集概述
数据集基本信息
- 数据集名称:VQ-VA World
- 任务类型:视觉问答-视觉回答(Visual Question–Visual Answering,VQ-VA)
- 核心功能:生成图像(而非文本)来回答视觉问题
数据规模与构建
- 数据量:约180万高质量交错的图像-文本样本
- 构建框架:基于智能代理管道的大规模定向数据构建
- 构建流程:
- 预处理阶段:对网络交错文档进行分类和过滤
- 代理管道:包含检索器、过滤器、指令生成器、重写器和推理器五个子模块
评估基准
IntelligentBench
- 样本数量:360个人工精选示例
- 评估维度:
- 世界知识:171个样本
- 设计知识:88个样本
- 推理能力:101个样本
- 构建流程:
- 文档审查:专家从约3000个分类交错网络文档中筛选最佳图像对
- 问题设计:专家针对每个图像对设计自由形式问题
- 专家交叉评审:每个候选项目需获得一致认可
其他评估基准
- RISEBench:基于推理的图像编辑基准
- KRIS-Bench:基于推理的图像编辑基准
- G-Edit-Benchmark-EN:标准图像编辑基准
- Img-Edit:标准图像编辑基准
性能表现
IntelligentBench 结果
- Ours模型:
- 世界知识:50.58
- 设计知识:57.95
- 推理能力:52.97
- 总体得分:53.06
- 对比模型:
- LightFusion(原始):7.78
- UniWorld-V1:1.94
- NanoBanana:81.67
- GPT-Image:82.64
资源发布
- 完整模型权重
- 数据集
- 构建管道
研究意义
通过发布完整的模型权重、数据集和管道,旨在促进VQ-VA领域的未来研究。




