ise-ice-lab/PSI_VQA
收藏资源简介:
PSI-VQA是一个视频问答基准数据集,基于PSI 2.0数据集构建,涵盖行人过马路场景的驾驶员视角行车记录仪视频。该数据集是AI City Challenge 2026 Track 3(驾驶员情境感知)的OOD测试集2。数据集包含四个互补任务:二值过马路问答(BCQ)、开放式问答(Open QA)、多项选择问答(MCQ)和时间定位(Temporal Localization),所有任务均统一遵循tao-vl-reason-v1.0模式(NVIDIA TAR Benchmark格式)。每个数据项将一个短视频片段与一个问题配对,模型需要返回结构化答案。训练集总共有1147个样本,测试集有328个样本。数据集用于评估模型在交通场景中的视觉推理和情境感知能力,仅限于学术和非商业研究使用。
PSI-VQA is a video question-answering benchmark built on the PSI 2.0 dataset, covering egocentric dashcam footage of pedestrian crossing scenarios. It serves as the OOD Test Set 2 for AI City Challenge 2026, Track 3: Driver Situation Awareness. The dataset spans four complementary tasks: Binary Crossing QA (BCQ), Open QA, Multiple Choice QA (MCQ), and Temporal Localization, all unified under the tao-vl-reason-v1.0 schema (NVIDIA TAR Benchmark format). Each item pairs a short video clip with a question, requiring the model to return a structured answer. The training set contains 1147 samples, and the test set has 328 samples. It is designed for evaluating visual reasoning and situational awareness in traffic contexts, restricted to academic and non-commercial research use.





