Molmo2-VideoCapQA
收藏资源简介:
# Molmo2-VideoCapQA Molmo2-VideoCapQA is a dataset of multiple-choice video QA that only requires visual content. It can be used to fine-tune vision-language models. Molmo2-VideoCapQA is part of the [Molmo2 dataset collection](https://huggingface.co/collections/allenai/molmo2-data) and was used to train the [Molmo2 family of models](https://huggingface.co/collections/allenai/molmo2). Quick links: - 📃 [Paper](https://allenai.org/papers/molmo2) - 🎥 [Blog with Videos](https://allenai.org/blog/molmo2) ## Data Format Videos are stored as Youtube video ID that will need to be downloaded separately. We provide a mapping from their IDs to the original YouTube URLs and public Google Cloud Storage URLs in `youtube_id_to_urls_mapping.json`. ### Video Download Videos are from YouTube. A mapping of video IDs to download URLs is provided in `youtube_id_to_urls_mapping.json`. The mapping contains more videos than this dataset needs — use the script below to download only the required ones. **Note:** Downloading from GCP URLs incurs GCP egress costs at the downloader's expense. ```python import json import datasets # Load the dataset to get the required video IDs ds = datasets.load_dataset("allenai/Molmo2-VideoCapQA", split="CapQA") needed_ids = set(ds["video_id"]) # Load the URL mapping and write needed GCP paths mapping = json.load(open("youtube_id_to_urls_mapping.json")) with open("files_to_download.txt", "w") as f: for vid_id in needed_ids: if vid_id in mapping: gcp_url = mapping[vid_id]["gcp_url"] gs_path = gcp_url.replace("https://storage.googleapis.com/", "gs://") f.write(gs_path + "\n") ``` Step 2: Download using gsutil: `cat files_to_download.txt | gsutil -m cp -I molmo2-youtube-cc/` The resulting directory structure should be: ``` molmo2-youtube-cc/ ├── youtube-cc-exist/{video_id}/{video_id}.{ext} ├── youtube-cc-kw/{video_id}/{video_id}.{ext} └── youtube-cc-temporal/{video_id}/{video_id}.{ext} ``` Alternatively, you can use the youtube_url field in the mapping to locate and download videos directly from YouTube. ## License This dataset is licensed under ODC-BY. A subset of videos from this dataset that are licensed as CC BY-4.0 may be downloaded from our Google Cloud Bucket via the URLs in `youtube_id_to_urls_mapping.json`. The dataset and videos are intended for research and educational use in accordance with Ai2’s [Responsible Use Guidelines](https://allenai.org/responsible-use). This dataset includes questions generated from GPT-4.1 and GPT-5, which are subject to OpenAI’s [Terms of Use](https://openai.com/policies/row-terms-of-use/).



