internlm/CapRL-Video-QA-20K
收藏资源简介:
CapRL-Video-QA-20K.jsonl是一个用于视频问答的数据集,作为CapRL++项目的一部分,旨在支持基于强化学习的视频描述生成任务。该数据集包含20,000个条目,每个条目指定了视频文件的相对路径,这些路径基于Hugging Face数据集lmms-lab/LLaVA-Video-178K的根目录。视频来源包括YouTube视频和学术视频(如Charades、NextQA、activitynet等),存储为压缩归档文件。数据集主要用于训练和评估CapRL-Video-4B模型,通过提供视频路径和问答对,帮助模型学习生成密集、时间戳标注的视频描述,以提升视频理解能力。数据集需要从指定源下载并按照特定目录结构组织,以确保路径正确性。
CapRL-Video-QA-20K.jsonl is a video question answering (VideoQA) dataset as part of the CapRL++ project, designed to support reinforcement learning-based video caption generation tasks. This dataset contains 20,000 entries, each specifying the relative path to a video file, with all paths based on the root directory of the Hugging Face dataset lmms-lab/LLaVA-Video-178K. The video sources cover YouTube videos and academic video datasets such as Charades, NextQA, ActivityNet, etc., and the videos are stored as compressed archives. The dataset is primarily intended for training and evaluating the CapRL-Video-4B model: by providing video paths and question-answer pairs, it helps the model learn to generate dense, timestamp-annotated video captions to enhance video understanding capabilities. The dataset needs to be downloaded from the specified source and organized following a specific directory structure to ensure correct path validity.




