GenS-Video150K
收藏资源简介:
GenS-Video150K是一个大规模合成的视频问答数据集,包含丰富的视频问题相关的帧注释。该数据集由北京大学、Salesforce Research和独立研究者共同构建,提供了大约20%的帧被标记为相关,并为每个相关帧分配了1到5级的细化置信度评分。数据集通过一个四阶段的管道生成,利用GPT-4o进行密集的视频帧字幕标注、构建视频问答对、扩展相关帧集合以及为相关帧打分。该数据集旨在帮助训练视频问答助手,更好地理解长时间视频内容。
GenS-Video150K is a large-scale synthetic video question answering (VideoQA) dataset with abundant frame annotations relevant to video questions. Developed jointly by Peking University, Salesforce Research, and independent researchers, this dataset marks approximately 20% of frames as relevant, and assigns a refined confidence score ranging from 1 to 5 to each relevant frame. The dataset is generated through a four-stage pipeline, which utilizes GPT-4o to conduct dense video frame captioning, build video question-answer pairs, expand the collection of relevant frames, and assign scores to relevant frames. This dataset is designed to facilitate the training of video QA assistants to better comprehend long-form video content.




