Candor-LR
收藏资源简介:
Candor-LR是一个面向自然对话场景的双人音频-视觉语音识别基准数据集,由都柏林圣三一大学Sigmedia研究组基于CANDOR语料库构建。该数据集汇集1,656场自发式视频会议对话,总计787,670条话语(约783.7小时),其中训练集占713.5小时,验证集10.1小时,测试集60.1小时,涵盖1,554位多样化说话人。创建过程采用定制化流水线:通过Speechmatics获得词级精细时间戳,实施说话人-视频映射与短语级分割(2-5秒),结合RetinaFace进行面部检测及嘴部ROI提取,并经过文本归一化与多重质量过滤,最终以说话人无重叠方式划分数据。该数据集旨在挑战现有AVSR模型在真实对话中的泛化能力,特别针对重叠语音、自发话轮转换与非脚本词汇等复杂条件,推动音频-视觉语音识别从广播化评估迈向更贴近实际应用的对话范式。
Candor-LR is a two-party audio-visual speech recognition (AVSR) benchmark dataset designed for natural conversational scenarios, developed by the Sigmedia Research Group at Trinity College Dublin based on the CANDOR corpus. This dataset contains 1,656 spontaneous video conference conversations, totaling 787,670 utterances (approximately 783.7 hours). The data is split into training, validation, and test sets with durations of 713.5 hours, 10.1 hours, and 60.1 hours respectively, covering 1,554 diverse speakers. The construction of this dataset follows a customized pipeline: word-level fine-grained timestamps are obtained through Speechmatics; speaker-video mapping and phrase-level segmentation (2–5 seconds per segment) are implemented; facial detection and mouth region of interest (ROI) extraction are performed using RetinaFace; followed by text normalization and multi-stage quality filtering. Finally, the dataset is split in a non-overlapping speaker-wise manner. This dataset aims to challenge the generalization capabilities of existing AVSR models in real-world conversational scenarios, particularly under complex conditions such as overlapping speech, spontaneous turn-taking, and unscripted vocabulary. It seeks to advance audio-visual speech recognition from broadcast-style evaluation toward more practical conversational paradigms that align with real-world applications.

- 1Candor-LR: A Dyadic Conversational Dataset for Audio-Visual Speech Recognition都柏林圣三一大学·工程学院·Sigmedia研究组 · 2026年



