OpenDriveLab/SparseVideoNav
收藏资源简介:
SparseVideoNav数据集专注于现实世界中的视觉语言导航任务,并结合稀疏未来视频生成技术。该数据集包含语言指令、RGB帧序列以及低级导航动作,其中每个发布的episode中的动作数量与RGB帧数量相匹配。数据集包含两个子集:BVN(超越视野导航)和IFN(指令跟随导航)。完整数据集总时长约为140小时,但由于地区政策限制,当前开源部分约为121.74小时。具体统计如下:BVN子集包含5,433个episodes、825,786个RGB帧、时长57.35小时,用于超越视野导航任务;IFN子集包含6,260个episodes、927,268个RGB帧、时长64.39小时,用于指令跟随导航任务。数据集以压缩tar分片格式存储图像,以避免在Hugging Face存储库中产生大量小文件。
SparseVideoNav studies real-world vision-language navigation with sparse future video generation. The datasets contain language instructions, RGB frame sequences, and low-level navigation actions. The number of actions matches the number of RGB frames for every released episode. This repository version contains the processed IFN and BVN subsets used by SparseVideoNav. The complete dataset contains about 140 hours; due to regional policy restrictions, the currently open-sourced portion is approximately 121.74 hours. The BVN subset has 5,433 episodes, 825,786 RGB frames, 57.35 hours duration, and is for Beyond-the-View Navigation; the IFN subset has 6,260 episodes, 927,268 RGB frames, 64.39 hours duration, and is for Instruction-Following Navigation. Images are stored in compressed tar shards to avoid hundreds of thousands of small files in the Hugging Face repository.




