EchoFoley-6k
收藏资源简介:
EchoFoley-6k是由字节跳动智能创作团队联合多所高校构建的大规模视频-音频标注数据集,包含6000个高质量视频-指令对和42000个细粒度声音事件标注。该数据集从VGGSound和PE Video Dataset中筛选运动明显的视频,通过LLM生成创意故事框架,再由专家标注员细化时间边界和声音属性,最终形成包含14种主题、平均时长11秒的视频样本。数据集创新性地采用符号化声音事件表示方法(时间戳、语义描述、音频属性),支持实例级、组级和视频级的三层控制,为视频配音生成任务提供精细化的时空对齐和属性控制基准。
EchoFoley-6k is a large-scale video-audio annotated dataset constructed by ByteDance's Intelligent Creation Team in collaboration with multiple universities. It contains 6,000 high-quality video-instruction pairs and 42,000 fine-grained sound event annotations. This dataset selects videos with prominent motions from VGGSound and PE Video Dataset, generates creative story frameworks via LLMs, then has expert annotators refine the temporal boundaries and sound attributes, ultimately resulting in video samples covering 14 themes with an average duration of 11 seconds. The dataset innovatively adopts a symbolic sound event representation method (including timestamps, semantic descriptions, and audio attributes), supporting three-level control at instance-level, group-level, and video-level, thereby providing a fine-grained spatio-temporal alignment and attribute control benchmark for video dubbing generation tasks.




