CMD-AM
收藏资源简介:
CMD-AM数据集是75部动画电影的集合,带有全面注释,包括角色边界框、真实的音频描述和句子级别的说话人分割标签。该数据集旨在支持研究动画电影中的角色识别和跟踪,并解决现有识别系统在处理动画电影方面的挑战。数据集的创建过程涉及自动构建音频-视觉角色库,包括角色外观库和角色语音库。数据集的应用领域包括为视觉障碍观众生成音频描述和为听力障碍观众生成角色感知字幕,从而显著提高动画内容的可访问性和叙事理解。
The CMD-AM dataset is a collection of 75 animated films with comprehensive annotations, including character bounding boxes, realistic audio descriptions, and sentence-level speaker diarization labels. This dataset is designed to support research on character recognition and tracking in animated films, and address the challenges faced by existing recognition systems when processing animated films. The creation of the dataset involves the automated construction of audio-visual character repositories, including character appearance libraries and character speech libraries. Application scenarios of this dataset include generating audio descriptions for visually impaired audiences and generating character-aware subtitles for hearing-impaired audiences, thereby significantly improving the accessibility of animated content and narrative comprehension.

- 1通过视觉几何组,牛津大学工程科学系,英国 · 2025年



