OneStory多镜头视频数据集
收藏资源简介:
该数据集由Meta AI与哥本哈根大学联合构建,包含约6万条高质量多镜头视频序列,专为长程叙事一致性建模而设计。数据内容聚焦人类中心活动,通过三阶段流程(镜头检测、两阶段标注、质量过滤)从原始视频中提取,每个镜头配备具有指代关系的渐进式文本描述。区别于传统全局脚本标注,采用镜头级参照性标注策略,确保叙事灵活性与真实拍摄场景相符,支持复杂场景下的跨镜头上下文建模。数据集主要应用于多镜头视频生成领域,旨在解决现有方法在长程叙事一致性和时空推理方面的局限性。
This dataset was co-developed by Meta AI and the University of Copenhagen, comprising approximately 60,000 high-quality multi-shot video sequences specifically designed for long-range narrative consistency modeling. Focusing on human-centric activities, the dataset is extracted from raw videos through a three-stage workflow consisting of shot detection, two-stage annotation, and quality filtering. Each shot is paired with progressive textual descriptions that feature coreferential relationships. Unlike conventional global script annotation, it adopts a shot-level referential annotation strategy, which ensures that narrative flexibility aligns with real-world filming scenarios and enables cross-shot contextual modeling in complex scenes. Primarily utilized in the domain of multi-shot video generation, this dataset aims to address the limitations of existing methods in long-range narrative consistency and spatio-temporal reasoning.




