PaveBench
收藏资源简介:
PaveBench是一个用于路面病害感知和交互式视觉-语言分析的大规模基准数据集,基于中国辽宁省高速公路检测车辆采集的真实世界图像构建。数据集包含视觉和多模态两个子集:视觉子集提供20,124张512×512分辨率的路面图像,支持图像分类、目标检测和语义分割任务;多模态子集PaveVQA包含32,160个问答对,涵盖单轮对话、多轮交互和专家校正三种类型。数据集包含六种路面病害类别,并特别设计了包含视觉混淆模式(如污渍、阴影和道路标记)的硬干扰子集以增强鲁棒性评估。PaveVQA问题设计围绕实际检测需求,包括存在验证、病害分类、定位、定量分析、严重性评估和维护建议等。数据集支持分类、目标检测、语义分割和视觉问答四项核心任务,旨在为路面领域的精确视觉感知和交互式多模态推理提供统一基础。
PaveBench is a large-scale benchmark dataset for pavement disease perception and interactive vision-language analysis, constructed using real-world images collected by highway detection vehicles in Liaoning Province, China. The dataset comprises two subsets: the visual subset and the multimodal subset. The visual subset includes 20,124 pavement images with a resolution of 512×512, supporting image classification, object detection, and semantic segmentation tasks. The multimodal subset, named PaveVQA, contains 32,160 question-answer pairs covering three types: single-turn dialogue, multi-turn interaction, and expert correction. The dataset encompasses six categories of pavement diseases, and a hard-disturbance subset specially designed with visual confusion patterns such as stains, shadows, and road markings is incorporated to facilitate robustness evaluation. The questions in PaveVQA are developed around actual detection requirements, covering existence verification, disease classification, localization, quantitative analysis, severity assessment, maintenance suggestions, and other related aspects. The dataset supports four core tasks including classification, object detection, semantic segmentation, and visual question answering, aiming to provide a unified foundation for accurate visual perception and interactive multimodal reasoning in the pavement engineering field.
PaveBench 数据集概述
数据集基本信息
- 数据集名称: PaveBench
- 许可证: CC BY-NC-SA 4.0
- 语言: 英语 (en)
- 数据规模: 10K < n < 100K
- 发布状态: 将在相关论文正式接受后,依据 CC BY-NC-SA 4.0 许可证公开发布。
核心任务与领域
- 任务类别: 视觉问答、图像分割、目标检测、图像分类。
- 研究领域: 计算机视觉、视觉-语言、视觉问答、图像分割、目标检测、图像分类、多模态学习、基准测试。
数据集内容与构成
PaveBench 是一个用于路面病害感知和交互式视觉-语言分析的大规模基准数据集,构建于真实世界的高速公路检测图像之上。
视觉感知子集 (Multi-Task Visual Perception)
- 图像数量: 20,124 张高分辨率路面图像。
- 图像尺寸: 512 × 512 像素。
- 支持任务: 图像分类、目标检测、语义分割。
- 图像类别: 纵向裂缝、横向裂缝、鳄鱼裂缝、修补、坑洞、负样本。
- 关键特性: 包含精心筛选的困难干扰项,如路面污渍、树影、道路标记等,用于鲁棒性评估。
多模态子集 (PaveVQA)
- 问答对总数: 32,160 对。
- 构成:
- 单轮查询: 10,050 对。
- 多轮交互: 20,100 对。
- 纠错对: 2,010 对。
- 问题覆盖范围: 存在性验证、病害分类、定位、定量分析、严重性评估、维护建议等。
数据来源与特点
- 数据来源: 中国辽宁省,使用配备高分辨率线扫描相机的高速公路检测车采集。
- 图像特点: 顶视正射路面视图,保留了病害模式的几何特性,支持可靠的下游量化分析。
- 标注特点: 为多项路面病害任务提供统一标注,旨在连接视觉感知与交互式视觉-语言分析。
基准测试与实验
- 视觉感知评估: 支持在统一基准下的分类、检测和分割任务。在检测和分割任务中,纵向裂缝和横向裂缝被合并为线性裂缝。
- 多模态VQA评估: LoRA 微调显著提升了视觉语言模型在路面特定问答上的性能。
- 智能体增强的VQA框架: 通过使用专门的视觉工具来锚定视觉语言模型的响应,提高了定量分析的可靠性。
引用信息
如需在您的工作中使用此数据集,请引用: bibtex @article{li2026pavebench, title={PaveBench: A Versatile Benchmark for Pavement Distress Perception and Interactive Vision-Language Analysis}, author={Li, Dexiang and Che, Zhenning and Zhang, Haijun and Zhou, Dongliang and Zhang, Zhao and Han, Yahong}, journal={arXiv preprint arXiv:2604.02804}, year={2026}, url={https://arxiv.org/abs/2604.02804} }




