遇见数据集

VLM3R-VSIBench-dataset

收藏
魔搭社区2026-08-01 更新2026-08-02 收录
官方服务:

资源简介:

# VLM-3R self-contained VSiBench training snapshot This directory is the canonical local training-data root. Training should use `data/vlm_3r_data` for both `--image_folder` and `--video_folder`. ## Snapshot contract - `vsibench/` contains the three files from the only public VSiBench release, Hugging Face commit `92fba1f605e3aab8de93e3e5ffb3848f271fdbc9`. - `scannet/videos/` contains 1,201 complete 640x480, 24 FPS videos exported frame-for-frame from ScanNet `.sens` RGB streams. - `scannetpp/videos/` contains the exact 856 `nvs_sem_train` scenes, sourced from ScanNet++ v2 `iphone/rgb.mkv`, resized to 640x480 with full frame count and source FPS preserved. - `arkitscenes/videos/` contains 348 raw Training-scene streams resized to 640x480 with full frame count and source FPS preserved. No manual rotation correction is applied, matching the authors' issue clarification. - Every dataset file is a regular local file with link count one. There are no symbolic links or hard links to the original shared datasets. - The loader uniformly samples 32 frames from each complete video and uses the real video duration to construct its time instruction. The paper and project README state that route planning has 4,225 rows. The only released `merged_qa_route_plan_train.json` has 4,104 rows; this snapshot pins the released file instead of synthesizing the unavailable 121 rows. ## Integrity verification `vsibench_snapshot_manifest.json` records video dimensions, FPS, frame counts, byte sizes, and SHA-256 hashes. Verify the complete snapshot without any raw source dataset: ```bash /vepfs-cnbje63de6fae220/shaolong/miniforge3/envs/vlm3r/bin/python \ scripts/VLM_3R/prepare_vsibench_train_data.py --verify-only ``` The optional `spatial_features/` directories are a derived cache, not source data. They may be regenerated from these videos and the pinned CUT3R weights.

提供机构:
maas
创建时间:
2026-07-27
二维码
社区交流群
二维码
科研交流群
商业服务