LibriheavyMix
收藏资源简介:
LibriheavyMix是一个由小米公司、腾讯AI实验室和香港中文大学联合创建的大规模单通道混响多说话者语音分离数据集,总时长达到20,000小时。该数据集基于Libriheavy构建,包含丰富的标点、大小写和文本上下文信息,旨在模拟真实世界的会议和鸡尾酒会场景。数据集的创建过程包括语音重叠模拟和混响引入,以生成更具挑战性的训练样本。LibriheavyMix主要应用于多说话者语音识别、语音分离和说话者日志,旨在解决在混响环境中识别“谁说了什么以及何时说”的难题。
LibriheavyMix is a large-scale single-channel reverberant multi-speaker speech separation dataset jointly created by Xiaomi Corporation, Tencent AI Lab, and The Chinese University of Hong Kong, with a total duration of 20,000 hours. Built upon Libriheavy, this dataset contains rich punctuation, capitalization, and textual context information, which is designed to simulate real-world meeting and cocktail party scenarios. The dataset creation process includes speech overlap simulation and reverberation injection to generate more challenging training samples. LibriheavyMix is mainly applied to multi-speaker speech recognition, speech separation and speaker diarization, aiming to solve the challenge of identifying "who spoke what and when" in reverberant environments.




