遇见数据集

Vript

收藏
魔搭社区2026-05-21 更新2024-05-15 收录
官方服务:

资源简介:

🎬 Vript: Refine Video Captioning into Video Scripting We construct a fine-grained video-text dataset with 12K annotated high-resolution videos (~400k clips). The annotation of this dataset is inspired by the video script. If we want to make a video, we have to first write a script to organize how to shoot the scenes in the videos. To shoot a scene, we need to decide the content, shot type (medium shot, close-up, etc), and how the camera moves (panning, tilting, etc). Therefore, we extend video captioning to video scripting by annotating the videos in the format of video scripts. Different from the previous video-text datasets, we densely annotate the entire videos without discarding any scenes and each scene has a caption with ~145 words. Besides the vision modality, we transcribe the voice-over into text and put it along with the video title to give more background information for annotating the videos.

Vript is a fine-grained video-text dataset comprising high-resolution videos (over 400k clips) paired with 12,000 annotations. The annotations of this dataset are inspired by video scripts. To produce a video, one must first draft a script to outline the shooting plan for each scene in the video. For shooting a single scene, decisions must be made regarding its content, shot type (e.g., medium shot, close-up) and camera movements (e.g., pan, tilt). Thus, drawing inspiration from the structure of video scripts, we annotate videos following the standard video script format. Unlike prior video-text datasets, we conduct dense annotation across the entire video without discarding any scene, with each scene accompanied by a descriptive caption of approximately 145 words. In addition to the visual modality, we also transcribe voiceovers into text and align them with the video captions to provide additional contextual information for the video annotations.

提供机构:
maas
创建时间:
2025-05-04
搜集汇总
数据集介绍
Vript 数据集图片
背景与挑战
背景概述
Vript是一个细粒度的英文视频-文本数据集,包含约12K个高分辨率视频(约400K个片段),标注格式采用视频脚本风格,涵盖镜头类型、摄像机运动和内容描述等详细信息。数据集提供完整视频、剪辑视频、字幕和元数据,分辨率达720p,并严格限制为学术研究使用。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务