V-QPP-Bench
收藏资源简介:
V-QPP-Bench是由吉林大学和密歇根州立大学团队构建的首个专注于视觉查询预处理(V-QPP)的多模态检索增强生成基准数据集。该数据集包含46,700条经过严格流程生成的缺陷视觉查询,涵盖几何畸变、质量退化、语义模糊等10类缺陷类型,数据源来自InfoSeek和ViQuAE等权威知识库VQA数据集。通过逆向工程方法对原始图像施加旋转、翻转、噪声等可控合成扰动构建查询-工具-真值三元组,支持图像到文本、图像到图像等5种MRAG检索范式评估,旨在解决现实场景中视觉查询缺陷导致的检索失败问题,推动鲁棒多模态系统的开发。
V-QPP-Bench is the first multimodal retrieval-augmented generation (MRAG) benchmark dataset focused on visual query preprocessing (V-QPP), constructed by teams from Jilin University and Michigan State University. This dataset contains 46,700 defect-containing visual queries generated via a rigorous processing workflow, covering 10 types of defects including geometric distortion, quality degradation, and semantic ambiguity. Its data sources are authoritative VQA datasets such as InfoSeek and ViQuAE. Query-tool-ground truth triplets are constructed by applying controllable synthetic perturbations such as rotation, flipping, and noise injection to original images through reverse engineering. It supports evaluation across 5 MRAG retrieval paradigms including image-to-text and image-to-image. This benchmark aims to address the retrieval failure issues caused by defective visual queries in real-world scenarios, and promote the development of robust multimodal systems.
- 1Fix Before Search: Benchmarking Agentic Query Visual Pre-processing in Multimodal Retrieval-augmented Generation吉林大学; 密歇根州立大学 · 2026年



