Audible623
收藏资源简介:
Audible623数据集是一个专门为可听动作时间定位任务设计的数据集,它从Kinetics和UCF101数据集中筛选出包含碰撞声音动作的视频,并对视频中的每个可听动作进行帧级标注。该数据集包含623个视频,平均每个视频250帧,为研究可听动作的时间定位提供了基础。数据集的创建旨在解决视频配音中可听动作的自动标记问题,以提高视频剪辑配音的效率。
The Audible623 Dataset is a specialized dataset designed for the task of audible action temporal localization. It selects videos containing actions accompanied by collision sounds from the Kinetics and UCF101 datasets, and performs frame-level annotations for each audible action in the videos. This dataset includes 623 videos, with an average of 250 frames per video, providing a foundational resource for research on temporal localization of audible actions. The dataset was developed to address the automatic annotation problem of audible actions in video dubbing, so as to improve the efficiency of video clip dubbing.
Audible623数据集概述
数据集基本信息
- 数据集名称: Audible623
- 关联研究: "Action Dubber: Timing Audible Actions via Inflectional Flow"
数据集用途
- 用于支持"Action Dubber"研究中关于通过屈折流定时可听动作的相关工作。
相关资源
- 包含数据集及配套代码。

- 1Action Dubber: Timing Audible Actions via Inflectional Flow新加坡管理大学计算与信息系统学院 · 2025年



