ov2_quickstart
收藏资源简介:
# OV2 Quickstart Quickstart bundle for **LLaVA-OneVision-2** (OV2). Contains everything needed to reproduce SFT training and run inference: packed SFT data, ready-to-use HF inference model, Megatron-Core checkpoint, and a Megatron training environment snapshot. **Total size:** ~374 GB across 329 files. --- ## Contents ### 1. `packed_mixed_sft_cap_v30s/` — 308 GB Packed mixed SFT (image + video + caption) dataset, sharded for distributed training via [Megatron-Energon](https://github.com/NVIDIA/Megatron-Energon). - **Format:** WebDataset shards (`.tar` + `.tar.idx`) - **Layout:** 4 nodes × 72 shards each ``` packed_mixed_sft_cap_v30s/ ├── dataset.yaml # Energon Metadataset config ├── node_a/webdataset/ # 77 GB — mixed_a-000000.tar … mixed_a-000035.tar (+ .idx) ├── node_b/webdataset/ # 78 GB — mixed_b-* ├── node_c/webdataset/ # 78 GB — mixed_c-* └── node_d/webdataset/ # 77 GB — mixed_d-* ``` - **Sample counts (from `dataset.yaml`):** ~508k samples per node, ~2.03M total - **Augmentation:** disabled (`augmentation: false`) **Use with Energon:** ```python from megatron.energon import get_train_dataset, WorkerConfig ds = get_train_dataset("packed_mixed_sft_cap_v30s/dataset.yaml", ...) ``` ### 2. `ov_encoder_p14m22_qwen3_hf/` — 8.9 GB HuggingFace-format **inference checkpoint** for LLaVA-OneVision-2 with Qwen3 LLM backbone. - **Architecture:** `LlavaOnevision2ForConditionalGeneration` - **LLM:** Qwen3-4B-Instruct-2507 (hidden_size=2560, intermediate_size=9728) - **Vision encoder:** patch-14, m22 variant - **Precision:** bfloat16 - **Custom modeling code** (trust_remote_code required): - `modeling_llava_onevision2.py` - `configuration_llava_onevision2.py` - `processing_llava_onevision2.py` - `codec_video_processing_llava_onevision2.py` - `video_processing_llava_onevision2.py` - **Demo script:** `demo_inference.py` ### 3. `ov_encoder_p14m22_qwen3_mcore_tp1pp1/` — 8.9 GB Equivalent **Megatron-Core checkpoint** of the same model, parallel layout `TP=1, PP=1`. Use this for continued training or fine-tuning in Megatron-LM / NeMo. ``` ov_encoder_p14m22_qwen3_mcore_tp1pp1/ ├── latest_checkpointed_iteration.txt └── release/ └── mp_rank_00/ └── model_optim_rng.pt ``` ### 4. `llava_megatron.26.05.tar` — 24 GB Frozen **training environment snapshot** (released 2025-05-26, hence `26.05`) containing the Megatron-LM fork, dependencies, and tooling used to produce the checkpoints in this repo. Provided as a tarball of an artifact directory (`blobs/sha256/...` content-addressed layout, 139 entries). **Extract:** ```bash tar -xf llava_megatron.26.05.tar ``` Use this to reproduce results bit-for-bit when external pip/git sources drift. --- ## Quickstart ```bash # Download just the inference model hf download lmms-lab-encoder/ov2_quickstart \ --repo-type dataset \ --include "ov_encoder_p14m22_qwen3_hf/*" \ --local-dir ./ov2 # Or pull everything (374 GB) hf download lmms-lab-encoder/ov2_quickstart \ --repo-type dataset \ --local-dir ./ov2 ``` --- ## File Manifest | Item | Size | Purpose | |---|---|---| | `packed_mixed_sft_cap_v30s/` | 308 GB | SFT training data (WebDataset, 4 nodes) | | `ov_encoder_p14m22_qwen3_hf/` | 8.9 GB | HF inference checkpoint | | `ov_encoder_p14m22_qwen3_mcore_tp1pp1/` | 8.9 GB | Megatron-Core training checkpoint | | `llava_megatron.26.05.tar` | 24 GB | Frozen training environment | | **Total** | **~374 GB** | |



