PhysicalAI-VANTAGE-Bench
收藏资源简介:
VANTAGE-BENCH 是首个专门用于评估固定基础设施摄像头捕获视频的视觉理解能力的公开基准数据集。该数据集涵盖三个现实世界领域——仓库、智能城市/智能交通系统(ITS)和智能空间,涉及六个时空视频理解任务,包括视频问答(VQA)、时间定位、密集视频字幕生成、事件验证、空间定位和时空跟踪。数据集仅用于评估目的。数据收集方法采用混合方式:人类采集、合成生成和自动化采集。视频数据来源包括供应商提供的素材、合成生成以及公开爬取的来源。标注方法同样采用混合方式:人工标注、合成标注和伪标注。数据集包含视频(mp4)和图像(jpg)格式,总存储量为42 GB。具体量化数据包括各任务的详细条目和帧数。数据集的所有评估均在服务器端进行,真实标注未公开发布。
VANTAGE-BENCH is the first publicly available benchmark dataset specifically designed to evaluate visual understanding capabilities of videos captured by fixed infrastructure cameras. This dataset covers three real-world domains: warehouses, smart cities/Intelligent Transportation Systems (ITS), and smart spaces, and involves six spatiotemporal video understanding tasks, including Video Question Answering (VQA), temporal localization, dense video captioning, event verification, spatial localization, and spatiotemporal tracking. This dataset is intended solely for evaluation purposes. It adopts a hybrid data collection method: human collection, synthetic generation, and automated collection. Video data sources include vendor-provided footage, synthetically generated content, and publicly crawled sources. The annotation method also uses a hybrid approach: manual annotation, synthetic annotation, and pseudo-labeling. The dataset contains videos in mp4 format and images in jpg format, with a total storage size of 42 GB. Specific quantitative data includes detailed entries and frame counts for each task. All evaluations of the dataset are conducted server-side, and the ground-truth annotations are not publicly released.




