S-VideoXum
收藏资源简介:
S-VideoXum数据集是一个扩展的VideoXum数据集,用于脚本驱动的视频摘要。该数据集包含了从VideoXum中提取的大量视频及其人类标注的摘要,并为每个视频的摘要添加了自然语言描述。这些描述是通过使用LLaVA-NeXT-7B大型多模态模型生成的。数据集的目的是支持脚本驱动的视频摘要研究,通过提供视频、摘要和摘要描述的三元组,可以训练模型生成适应不同用户需求的视频摘要。S-VideoXum数据集拥有超过11908个视频条目,覆盖了多个领域,视频长度最长可达12.5分钟。该数据集的创建旨在解决现有视频摘要方法无法根据用户特定需求生成摘要的问题,通过脚本驱动的摘要方法,可以生成更丰富、更符合用户需求的视频摘要。
The S-VideoXum dataset is an extended version of the VideoXum dataset, tailored for script-driven video summarization. This dataset contains a large number of videos extracted from VideoXum along with their human-annotated summaries, and adds natural language descriptions for each video's summary. These descriptions are generated using the LLaVA-NeXT-7B large multimodal model. The dataset aims to support research on script-driven video summarization: by providing triplets of videos, summaries, and summary descriptions, it enables the training of models to generate video summaries tailored to diverse user requirements. The S-VideoXum dataset comprises over 11,908 video entries, covering multiple domains, with the maximum video duration reaching up to 12.5 minutes. This dataset was developed to address the limitation of existing video summarization methods that fail to generate summaries tailored to specific user needs; script-driven summarization approaches can produce more abundant and user-aligned video summaries.

- 1SD-VSum: A Method and Dataset for Script-Driven Video Summarization希腊塞萨洛尼基CERTH-ITI · 2025年



