MultiVSR2LRS3 (MV2LRS3)
收藏资源简介:
MultiVSR2LRS3是由帝国理工学院和都柏林三一学院联合构建的音频-视觉语音识别评估数据集,旨在严格评估现有模型的真实泛化性能。该数据集从大规模MultiVSR语料库的英语验证集中精心采样,包含与经典LRS3测试集在声学、视觉和人口统计学等七个关键维度上严格匹配的分布,数据规模与LRS3测试集相当(约0.9小时视频)。其构建过程采用多维k近邻匹配策略,通过加权特征向量实现分布对齐,确保评估的严谨性。该数据集主要应用于揭示音频-视觉语音识别模型在分布匹配条件下的泛化缺陷,为解决当前模型过度适应特定基准、缺乏鲁棒性的核心问题提供关键评估工具。
MultiVSR2LRS3 is an audio-visual speech recognition evaluation dataset jointly constructed by Imperial College London and Trinity College Dublin, aiming to rigorously evaluate the true generalization performance of existing models. It is meticulously sampled from the English validation subset of the large-scale MultiVSR corpus, and its data distribution strictly matches that of the classic LRS3 test set across seven key dimensions including acoustics, vision and demographics. The scale of this dataset is comparable to that of the LRS3 test set, containing approximately 0.9 hours of video. The dataset is built using a multi-dimensional k-nearest neighbor matching strategy, which achieves distribution alignment via weighted feature vectors to ensure the rigor of the evaluation. This dataset is primarily used to reveal the generalization flaws of audio-visual speech recognition models under distribution-matched conditions, providing a critical evaluation tool to address the core issue that current models overly adapt to specific benchmarks and lack robustness.
数据集名称: MultiVSR2LRS3 (MV2LRS3)
所属论文: Assessing True Generalisability of Audio-Visual Speech Recognisers(Interspeech 2026 长文)
核心用途: 评估语音识别模型真实泛化能力的音视频语音数据集。
数据来源与构建: 从 MultiVSR 数据集中严格筛选子样本,使其分布与 LRS3 测试集分布完全一致。
数据集下载与结构:
- 下载地址:https://drive.google.com/file/d/1jmH1srv_27M19cC0JblvV9CuuW9WxjeS/view?usp=drive_link
- 目录结构:
data/:原始媒体文件,包含裁剪嘴部区域的视频(mp4格式)和对应音频(wav格式)。manifest/:评估子集文件,包含6个TSV文件:subset_v1.tsv至subset_v5.tsv:5个不同随机选择的受控子集。10x.tsv:10倍缩放集。
- TSV 文件内容:包含转录文本及所有元数据(时长、年龄、性别、表观肤色、头部姿态、信噪比、语速)。
评估方法: 推荐在全部5个子集上运行推理,并报告平均性能指标和标准差。
与 LRS3 测试集的对比: 提供了核心统计对比图,证明 MV2LRS3 的分布与 LRS3 测试集严格一致。
伦理声明: 人口统计元数据通过自动化工具生成,属于算法推断,非自我报告身份,研究者应谨慎使用这些标签。




