MAD (Movie Audio Descriptions)
收藏资源简介:
MAD数据集是由阿卜杜拉国王科技大学创建的一个大规模视频语言接地基准,包含超过384,000个自然语言句子,这些句子与超过1,200小时的视频内容相对应。该数据集通过爬取和校准主流电影的音频描述来构建,旨在减少现有视频语言接地数据集的偏见。MAD数据集的收集策略使得视频语言接地任务更具挑战性,要求在长达三小时的多样的长格式视频中准确接地短时间(通常为几秒)的时刻。该数据集广泛应用于智能视频搜索、视频编辑和帮助记忆障碍患者等领域,为解决视频语言接地问题提供了丰富的资源。
The MAD Dataset is a large-scale video-language grounding benchmark developed by King Abdullah University of Science and Technology. It contains over 384,000 natural language sentences aligned with more than 1,200 hours of video content. Constructed by crawling and calibrating audio descriptions of mainstream films, this dataset is designed to mitigate biases inherent in existing video-language grounding datasets. The collection strategy employed for the MAD Dataset elevates the difficulty of video-language grounding tasks, requiring accurate grounding of short-duration (typically several seconds) moments within diverse long-form videos that can span up to three hours. This dataset has been widely adopted in applications including intelligent video search, video editing, and assisting patients with memory impairments, serving as a rich resource for advancing solutions to video-language grounding tasks.




