Rank2Tell
收藏资源简介:
Rank2Tell是由本田美国研究所和斯坦福大学合作开发的多模态数据集,专注于城市交通场景中的视觉场景理解和重要性排序。该数据集包含116个视频片段,每个约20秒,聚焦于交叉路口的复杂交通情况。数据集通过视觉问答(VQA)方法,提供了密集的语义、空间、时间和关系属性标注,以及自然语言解释,以增强自动驾驶系统的透明度和可解释性。此外,数据集还引入了联合模型,用于同时预测重要性级别和生成自然语言标题,为安全关键应用提供了新的研究资源。
Rank2Tell is a multimodal dataset jointly developed by Honda Research Institute USA and Stanford University, focusing on visual scene understanding and importance ranking in urban traffic scenarios. This dataset contains 116 video clips, each approximately 20 seconds long, focusing on complex traffic situations at intersections. Through visual question answering (VQA) methods, it provides dense semantic, spatial, temporal and relational attribute annotations as well as natural language explanations, aiming to enhance the transparency and interpretability of autonomous driving systems. Additionally, the dataset introduces a joint model for simultaneously predicting importance levels and generating natural language captions, offering a novel research resource for safety-critical applications.




