CHILDES-Aligned
收藏资源简介:
CHILDES-Aligned 是由伊利诺伊大学厄巴纳-香槟分校等机构联合创建的一个经过精心整理的儿童语音数据集,旨在解决原始CHILDES语料库中话语级时间戳噪声大、不完整或与音频未对齐的问题。该数据集包含413小时的自然儿童-成人互动长时录音,并提供了经过BEACON框架校正的精确话语级时间戳,同时从中提取了一个283小时的高质量子集用于语音识别训练。数据创建过程采用了多模型时间戳集成技术,通过聚合多个现成的ASR模型的预测结果,并采用共识投票策略来生成可靠的时间边界。该数据集主要应用于儿童语音识别、语言发展研究以及语音模型的训练与评估领域,旨在提升模型在真实儿童语音场景下的鲁棒性和准确性。
CHILDES-Aligned is a meticulously curated child speech dataset jointly created by the University of Illinois Urbana-Champaign and other institutions, aiming to resolve the issues of excessive utterance-level timestamp noise, incompleteness, and audio misalignment in the original CHILDES corpus. This dataset contains 413 hours of long-duration natural child-adult interactive recordings, and provides precise utterance-level timestamps corrected via the BEACON framework. Additionally, a 283-hour high-quality subset is extracted from it for speech recognition training. The data creation process adopts multi-model timestamp integration technology, which aggregates predictions from multiple off-the-shelf ASR models and employs a consensus voting strategy to generate reliable temporal boundaries. This dataset is primarily applied in the fields of child speech recognition, language development research, as well as speech model training and evaluation, with the goal of improving the robustness and accuracy of models in real-world child speech scenarios.

- 1CHILDES-Aligned: A Curated Children's Speech Dataset via Multi-Model Timestamp Ensembling伊利诺伊大学厄巴纳: 香槟分校; 香港中文大学·深圳; IBM研究院; 布法罗大学 · 2026年



