672小时多人会议多通道采集语音数据
收藏资源简介:
672小时多人会议多通道采集语音数据,内容覆盖3-6人规模的会议场景,在多种会议室环境中采集,反映了真实会议中的互动情境。此数据集标注了文本内容、说话人身份、性别及位置等多种属性,准确性高(句准确率≥97%),易用性强,为语音识别及声纹识别相关研究与应用提供了高质量资源,经多家AI公司验证:有助于提升模型在复杂会议场景下的鲁棒性。我们严格遵循数据保护法规和隐私规定,确保数据采集、存储和使用的过程中维护用户的隐私和合法权益,所有数据均遵循GDPR, CCPA, PIPL。
This is a 672-hour multi-channel speech dataset collected from multi-person meetings. It covers meeting scenarios with 3 to 6 participants, and was collected across various conference room environments, thus reflecting real interaction scenarios in actual meetings. The dataset is annotated with multiple attributes including text content, speaker identity, gender and location, featuring high accuracy (sentence-level accuracy ≥ 97%) and excellent usability. It provides high-quality resources for research and applications related to speech recognition and speaker verification, and has been validated by multiple AI companies that it can effectively improve the robustness of models in complex meeting scenarios. We strictly comply with data protection regulations and privacy rules, ensuring that users' privacy and legitimate rights and interests are protected throughout the data collection, storage and utilization processes. All data complies with GDPR, CCPA and PIPL.




