LRS3-TED
收藏资源简介:
LRS3-TED数据集是由牛津大学视觉几何组创建的大规模多模态数据集,主要用于视觉和音频-视觉语音识别。该数据集包含超过400小时的TED和TEDx视频中的面部轨迹,以及相应的字幕和单词对齐边界。数据集内容丰富,包含5594个视频,每个视频的面部轨迹以224×224分辨率和25 fps帧率提供。数据集的创建过程涉及多阶段自动化管道,用于生成大规模的音频-视觉语音识别数据。LRS3-TED数据集广泛应用于唇读、音频-视觉语音识别等领域,旨在解决缺乏大规模公共基准数据集的问题。
The LRS3-TED dataset is a large-scale multimodal dataset developed by the Visual Geometry Group at the University of Oxford, primarily designed for visual and audio-visual speech recognition tasks. It encompasses facial trajectories extracted from more than 400 hours of TED and TEDx videos, along with paired subtitles and word-aligned boundaries. The dataset consists of 5,594 videos in total, with facial trajectories for each video offered at a resolution of 224×224 and a frame rate of 25 fps, boasting rich and diverse content. The development of this dataset employs a multi-stage automated pipeline to generate large-scale audio-visual speech recognition data. Widely applied in domains such as lip reading, audio-visual speech recognition and other relevant fields, the LRS3-TED dataset is intended to address the shortage of large-scale public benchmark datasets.




