RoomReader-AV
收藏资源简介:
RoomReader-AV是都柏林圣三一大学推出的多条件音视频语音识别基准,旨在评估模型在非广播场景下的泛化能力。该数据集包含10324条自发多人在线会议语音片段,涵盖118位发言人的真实对话,总时长约6.5小时。数据收集自30场Zoom辅导课,经人工校正转写并保留语气词与填充停顿,预处理后提供标准化的音视频ROI裁剪与统一格式。该数据集主要应用于音视频语音识别系统的鲁棒性评估,尤其针对摄像角度、说话人发音风格及自发对话等挑战性条件,以揭示当前模型在受控广播数据之外的性能退化问题。
RoomReader-AV is a multi-condition audio-visual speech recognition benchmark developed by Trinity College Dublin, aimed at evaluating the generalization ability of models in non-broadcast scenarios. This dataset contains 10,324 spontaneous multi-party online meeting speech segments, covering real dialogues from 118 speakers, with a total duration of approximately 6.5 hours. The data was collected from 30 Zoom tutoring sessions, with manually corrected transcriptions that retain filler words and filled pauses, and standardized audio-visual ROI cropping and unified formats are provided after preprocessing. This dataset is mainly applied to the robustness evaluation of audio-visual speech recognition systems, especially under challenging conditions such as camera angles, speakers' pronunciation styles and spontaneous dialogues, to reveal the performance degradation of current models beyond controlled broadcast datasets.
Lipreading Data Guide 数据集概述
项目简介
一个用于将主流唇读和音视频语音识别数据集预处理为统一标准化格式的综合工具包。自动化预处理步骤包括人脸跟踪、嘴部 ROI 提取和转录对齐,以简化模型训练工作流。与 AV-HuBERT 和 Auto-AVSR 兼容。
支持的数据集
| 数据集 | 状态 | 描述 | 规模 | 主要特征 |
|---|---|---|---|---|
| LRS2 | ✅ Ready | Lip Reading Sentences 2 | ~140k utterances | BBC 广播,野外场景 |
| LRS3 | ✅ Ready | Lip Reading Sentences 3 | ~150k utterances | TED 演讲,高质量 |
| LRS_Combined | ✅ Ready | LRS2 + LRS3 合并 | ~290k utterances | 统一语料库 |
| TCD-TIMIT | ✅ Ready | TCD-TIMIT 音视频语料库 | ~27k utterances | 高清视频,受控环境,2 个摄像机角度 |
| GRID | ✅ Ready | 音视频语音语料库 | ~34k utterances | 受控词汇,34 位说话人 |
| LombardGrid | ✅ Ready | Lombard 效应语音 | ~5.4k utterances | 噪声条件,54 位说话人 |
| RoomReader | ✅ Ready | 多方对话 | ~322 videos | 在线会议,118 位参与者 |
| Candor | ✅ Ready | 自然对话 | ~1,656 conversations | 无脚本双人对话 |
| VoxCeleb2 | 🔄 Planned | 说话人识别数据集 | ~1M utterances | 计划使用 Whisper 转录流程 |
| WildVSR | ✅ Ready | Wild VSR 测试集 | Test set | 泛化基准 |
| AVCocktail | ✅ Ready | 鸡尾酒会语音 | Challenge dataset | 多说话人,噪声 |
| Muavic | ⚠️ Experimental | 多语言音视频 | 9 languages | 语音识别 + 翻译 |
| MultiVSR | 🔄 Planned | 大规模多语言 VSR | ~1,400 hours | 20+ 种语言 |
状态说明
- ✅ Ready:已完整测试并可投入生产
- ⚠️ Experimental:可用但未完全测试
- 🔄 Planned:计划未来支持(可能提供下载说明)
实用工具
- Phones:音素转换和映射工具,用于音素级模型训练
- webData:将 Auto-AVSR 预处理数据转换为 WebDataset 和 Hugging Face 兼容格式的工具
快速开始
每个数据集文件夹包含全面的文档,提供数据获取、预处理和准备的分步说明。导航至特定数据集目录可查看详细设置指南和处理工作流。
路线图
计划通过增加对多语言和多方对话数据集的支持,将工具包扩展到以英语为中心的基准之外,包括 MultiVSR、MARC、MISP、MLD-VC、CI-AVSR、RUSAVIC、KMSAV、VISPER、Friends-MMC、AVSD、HAVRUS、ViCocktail、OLKAVS、F2F-JF 和 Seamless Interaction。这些数据集可有效转换为统一格式,使其可用于 AVSR 和 VSR 研究,并实现跨多种语言和对话场景的更全面、标准化的基准测试。
社区贡献
该工具包设计为社区驱动项目而非封闭框架。欢迎贡献以增加对新数据集的支持、改进数据集转换方案或引入标准化评估协议。目标是构建统一且可扩展的 AVSR 研究基础设施,实现跨多样化数据集的一致训练和评估。
数据集参考文献
- LRS2:Afouras, T., Chung, J. S., Senior, A., Vinyals, O., & Zisserman, A. (2018). Deep Audio-Visual Speech Recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence. https://www.robots.ox.ac.uk/~vgg/data/lip_reading/lrs2.html
- LRS3:Afouras, T., Chung, J. S., & Zisserman, A. (2018). LRS3-TED: a large-scale dataset for visual speech recognition. arXiv preprint arXiv:1809.00496. https://www.robots.ox.ac.uk/~vgg/data/lip_reading/lrs3.html
- TCD-TIMIT:Harte, N., & Gillen, E. (2015). TCD-TIMIT: An audio-visual corpus of continuous speech. IEEE Transactions on Multimedia, 17(5), 603-615. https://sigmedia.tcd.ie/TCDTIMIT/
- WildVSR:Djilali, Y. A. D., Narayan, S., LeBihan, E., Boussaid, H., Almazrouei, E., & Debbah, M. (2024). Do VSR Models Generalize Beyond LRS3? Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 6635-6644. https://github.com/YasserdahouML/VSR_test_set
- VoxCeleb2:Chung, J. S., Nagrani, A., & Zisserman, A. (2018). VoxCeleb2: Deep Speaker Recognition. Interspeech 2018. https://www.robots.ox.ac.uk/~vgg/data/voxceleb/vox2.html
- AVCocktail & AVYT:Nguyen, T.-B., Pham, N.-Q., Waibel, A. (2025). Cocktail-Party Audio-Visual Speech Recognition. Proc. Interspeech 2025, 1828-1832. https://arxiv.org/abs/2506.02178
- MuAViC:Anwar, A., Shi, B., Goswami, V., Hsu, W. N., Pino, J., & Wang, C. (2023). MuAViC: A Multilingual Audio-Visual Corpus for Robust Speech Recognition and Robust Speech-to-Text Translation. arXiv preprint arXiv:2303.00628. https://github.com/facebookresearch/muavic
- GRID:Cooke, M., Barker, J., Cunningham, S., & Shao, X. (2006). An audio-visual corpus for speech perception and automatic speech recognition. The Journal of the Acoustical Society of America, 120(5), 2421-2424. https://zenodo.org/records/3625687
- Lombard GRID:Alghamdi, N., Maddock, S., Marxer, R., Barker, J., & Brown, G. J. (2018). A corpus of audio-visual Lombard speech with frontal and profile views. The Journal of the Acoustical Society of America, 143(6), EL523-EL529. https://zenodo.org/records/3228148
- RoomReader:Reverdy, J., OConnor Russell, S., Duquenne, L., Garaialde, D., Cowan, B. R., & Harte, N. (2022). RoomReader: A Multimodal Corpus of Online Multiparty Conversational Interactions. Proceedings of the Thirteenth Language Resources and Evaluation Conference, 2517-2527. https://aclanthology.org/2022.lrec-1.268/
- MultiVSR:Prajwal, K. R., Hegde, S., & Zisserman, A. (2025). Scaling Multilingual Visual Speech Recognition. ICASSP 2025 - IEEE International Conference on Acoustics, Speech and Signal Processing, 1-5. https://github.com/Sindhu-Hegde/multivsr
- Candor:Reece, A., Cooney, G., Bull, P., & Chung, C. (2023). The CANDOR corpus: Insights from a large multimodal dataset of naturalistic conversation. Science Advances, 9, eadf3197. https://www.science.org/doi/10.1126/sciadv.adf3197 | https://candor.usc.edu/
代码库参考文献
AV-HuBERT
Shi, B., Hsu, W.-N., Lakhotia, K., & Mohamed, A. (2022). Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction (AV-HuBERT). 论文:https://arxiv.org/abs/2201.02184 GitHub:https://github.com/facebookresearch/av_hubert
Auto-AVSR
Ma, P., Haliassos, A., Fernandez-Lopez, A., Chen, H., Petridis, S., & Pantic, M. (2023). Auto-AVSR: Audio-Visual Speech Recognition with Automatic Labels. 论文:https://arxiv.org/abs/2303.14307 GitHub:https://github.com/mpc001/auto_avsr




