Reframed
收藏资源简介:
Reframed是由爱丁堡大学信息学院构建的高质量音频描述生成数据集,旨在解决电影中何时描述以及描述什么内容的联合决策问题。该数据集包含来自206部电影的2,023个视频片段,涵盖3,302个场景,并提供了专业的美式与英式音频描述转录、对话字幕、听觉障碍者字幕以及对齐的剧本等多模态数据。创建过程中,训练数据通过自动语音识别和自监督对齐流水线生成,而评估数据则采用专业人工转录与帧级时间编码,确保高精度。该数据集主要应用于电影音频描述自动化研究,旨在帮助视障人士理解电影叙事,同时为多模态大语言模型在视频理解与生成任务中提供新的基准与挑战。
Reframed is a high-quality audio description generation dataset developed by the School of Informatics, University of Edinburgh, aiming to address the joint decision-making problem of when and what to describe in films. This dataset comprises 2,023 video clips from 206 films, spanning 3,302 scenes, and provides multi-modal data including professional American and British audio description transcripts, dialogue subtitles, subtitles for the hard of hearing, and aligned screenplays. During its development, the training data was generated through an automatic speech recognition and self-supervised alignment pipeline, while the evaluation data was produced via professional manual transcription and frame-level time coding to ensure high accuracy. This dataset is primarily applied to research on automated audio description for films, aiming to assist visually impaired people in understanding film narratives, while also providing new benchmarks and challenges for multimodal large language models in video understanding and generation tasks.




