AVA-ActiveSpeaker
收藏资源简介:
AVA-ActiveSpeaker数据集是由谷歌人工智能感知部门创建的一个大规模、多样化的音频-视觉数据集,旨在支持主动说话人检测任务。该数据集包含约3.65百万个标注帧,覆盖约38.5小时的人脸轨迹视频及其音频。数据集中的每个人脸实例都被标记为说话或不说话,以及语音是否可听见。数据集的构建过程涉及视频选择、标签词汇定义、人脸轨迹检测和人工标注。AVA-ActiveSpeaker数据集的应用领域广泛,包括说话人分割、视频会议重定向、语音增强和人与机器人交互等,旨在解决视频分析中的核心问题,如识别视频中哪个可见人物正在说话。
The AVA-ActiveSpeaker dataset is a large-scale, diverse audio-visual dataset developed by Google's AI Perception Department, which is designed to support active speaker detection tasks. It encompasses approximately 3.65 million annotated frames, spanning roughly 38.5 hours of videos with tracked face trajectories and their corresponding audio streams. Every individual face instance within the dataset is annotated with two pieces of information: whether the person is speaking, and whether the speech is audible. The development of the AVA-ActiveSpeaker dataset encompasses several key stages, including video selection, definition of labeling vocabulary, face trajectory detection, and manual annotation. This dataset has broad application prospects across multiple domains, such as speaker diarization, video conference redirection, speech enhancement, and human-robot interaction, among others. Its core goal is to address fundamental challenges in video analysis, such as identifying which visible individual in a video is currently speaking.




