CoReVLA
收藏资源简介:
CoReVLA数据集是一个用于端到端自动驾驶框架的双阶段数据集,旨在提高在长尾场景中的性能。数据集通过数据收集和行为细化两个阶段进行构建。首先,模型在多个开源驾驶问答数据集上进行监督微调,以获取对驾驶场景的基础理解。然后,CoReVLA在CAVE模拟平台中部署,从实时交互中收集司机接管数据。每个接管都表示CoReVLA无法可靠处理的长期场景。最后,通过直接偏好优化(DPO)对模型进行细化,使其能够直接从人类偏好中学习,从而避免由手动设计的奖励造成的奖励黑客攻击。
The CoReVLA dataset is a two-stage dataset tailored for end-to-end autonomous driving frameworks, aiming to improve performance in long-tail driving scenarios. It is constructed through two phases: data collection and behavior refinement. First, the model undergoes supervised fine-tuning on multiple open-source driving question-answering datasets to gain a foundational understanding of driving scenarios. Subsequently, CoReVLA is deployed on the CAVE simulation platform to collect human driver takeover data from real-time interactions. Each takeover represents a long-tail scenario that CoReVLA cannot handle reliably. Finally, the model is refined via Direct Preference Optimization (DPO), allowing it to learn directly from human preferences and thereby avoid reward hacking induced by manually designed reward functions.

- 1通过同济大学交通运输学院 · 2025年



