sayedpedramhaeri/VLA4CoDrive
收藏资源简介:
VLA4CoDrive是一个大规模协作视觉-语言-动作(VLA)数据集,专为支持多车辆协作下的自动驾驶而设计。该数据集包含同步的多车辆感知数据,覆盖多样化的驾驶环境,提供多视角视觉流、上下文文本注释(包括标题、上下文、描述和推理)以及未来轨迹动作,用于训练和评估VLA驾驶模型。数据集具有协作多车辆设置、多视角多模态感知(如RGB摄像头、LiDAR、语义LiDAR、光流、GNSS和IMU)、结构化视觉-语言基础、动作与轨迹监督(包括低层控制和30步未来轨迹)、受控多样性(8个CARLA城镇和8种天气条件)以及大规模数据(约1000万视觉样本、15万语言注释、100万动作记录和300-360小时驾驶数据),并采用标准注释格式(如COCO、PASCAL VOC、KITTI 2D和KITTI 3D)。
VLA4CoDrive is a large-scale cooperative Vision–Language–Action (VLA) dataset designed to support autonomous driving under multi-vehicle cooperation. It features synchronized multi-vehicle sensing across diverse driving environments, providing multi-view visual streams, contextual text annotations including caption, context, description, and reasoning, and future trajectory actions for training and evaluating VLA driving models. The dataset includes a cooperative multi-vehicle setup, multi-view and multi-modal perception (e.g., RGB front, rear, left, and right cameras, LiDAR, semantic LiDAR, optical flow, GNSS, and IMU), structured vision–language grounding, action and trajectory supervision (including low-level controls and 30-step future trajectories), controlled diversity (8 CARLA towns and 8 weather conditions), and large-scale data (approximately 10M vision samples, 150K language annotations, 1M action records, and 300–360 hours of driving data), with standard annotation formats such as COCO, PASCAL VOC, KITTI 2D, and KITTI 3D.



