遇见数据集

Molmo2-ER-VSI-590K

收藏
魔搭社区2026-07-13 更新2026-07-15 收录
官方服务:

资源简介:

# Molmo2-ER · nyu-visionx/VSI-590K 590K spatial QA samples (image+video) propagated from 3D ground truth and CV pseudo-labels. This is a re-hosted, **loader-ready subset** of the upstream dataset, used to train [`allenai/Molmo2-ER-4B`](https://huggingface.co/allenai/Molmo2-ER-4B). Files mirror the upstream layout; nothing in the data has been modified. ## Upstream source - **Original dataset:** [nyu-visionx/VSI-590K](https://huggingface.co/datasets/nyu-visionx/VSI-590K) - **Paper:** *Cambrian-S: Towards Spatial Supersensing in Video* ([arXiv:2511.04670](https://arxiv.org/abs/2511.04670)) - **License:** `apache-2.0` (inherits from upstream) If you use this data, please cite the original authors: ```bibtex @misc{yang2025cambriansspatialsupersensingvideo, title={Cambrian-S: Towards Spatial Supersensing in Video}, author={Shusheng Yang and Jihan Yang and Pinzhi Huang and others}, year={2025}, eprint={2511.04670}, archivePrefix={arXiv} } ``` ## Extracting before training This release ships archives. Extract them in-place before pointing `SPATIAL_DATA_HOME` at this directory: ```bash for f in *.tar.gz; do tar -xzf $f; done ``` ## Usage in Molmo2-ER See the [`allenai/molmo2`](https://github.com/allenai/molmo2) repository for the data loader and training recipe. The relevant loader class for this dataset lives in `olmo/data/spatial_datasets.py`.

提供机构:
maas
创建时间:
2026-05-06
二维码
社区交流群
二维码
科研交流群
商业服务