遇见数据集

LLaVA-OneVision-2-Data

收藏
魔搭社区2026-08-28 更新2026-08-30 收录
官方服务:

资源简介:

# LLaVA-OneVision-2-Data Training data for the LLaVA-OneVision-2 multimodal model family, covering large-scale video and spatial reasoning corpora used in mid-training. ## Dataset Composition | Subset | Format | Description | |---|---|---| | `mid_training_video/60s_rest/` | WebDataset (`.tar`) | 10,809 shards of ~60s video clips | | `mid_training_video/caption_v0/split_30s.jsonl` | JSONL | Captions for 30-second video clips | | `mid_training_video/caption_v0/split_60s.jsonl` | JSONL | Captions for 60-second video clips | | `mid_training_video/caption_v0/split_180s.jsonl` | JSONL | Captions for 180-second video clips | | `mid_training_video/caption_v0/split_gt10min.jsonl` | JSONL | Captions for >10-minute video clips | | `spatial/` | WebDataset (`.tar`) | 84 shards of spatial reasoning data (refcoco, visual genome, pointing, 3D, etc.) | | `mid_training_video/mapping/mapping_{5s,10s,30s,60s,180s,gt10min}.csv` | CSV | Maps each video clip's `dst_path` to its source `youtube_id` and `[start_time, end_time]` window | ## Preview Configs The `viewer_*` configs above expose small Parquet samples so the Hugging Face Dataset Viewer can render the data directly in the browser: - **`viewer_caption_30s`** — 5 caption samples from 30-second clips - **`viewer_caption_60s`** — 5 caption samples from 60-second clips - **`viewer_caption_180s`** — 3 caption samples from 180-second clips - **`viewer_caption_gt10min`** — 1 caption sample from >10-minute clips - **`viewer_spatial`** — 10 spatial-reasoning samples with embedded thumbnail images, mixed across tasks (refcoco, visual genome, pointing, ca1m, osd, crosspoint, erqa, roborefer) These previews are intended for **schema inspection only**. For training, use the full `mid_training_video/` and `spatial/` shards.

提供机构:
maas
创建时间:
2026-08-21
二维码
社区交流群
二维码
科研交流群
商业服务