M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset
收藏资源简介:
M3SD数据集是一个多模态、多场景和多语言的说话人分割数据集,旨在解决现有数据集规模不足和深度学习模型泛化能力差的问题。该数据集来源于真实网络视频,包含了770+小时的音频和视频数据,涵盖了访谈、线上/线下会议、演讲、辩论、日常对话等多种场景,并支持中、英、日等多种语言。数据集的创建过程采用了自动化方法,结合音频和视频信息生成更准确的伪标签,并通过预训练的说话人分割模型进行迭代训练。该数据集的发布为研究说话人分割技术提供了新的数据资源,有助于提升模型的泛化能力和适应不同场景的能力。
The M3SD dataset is a multimodal, multi-scenario, multilingual speaker diarization dataset aimed at addressing the issues of insufficient scale of existing datasets and poor generalization ability of deep learning models. Derived from real-world web videos, this dataset contains over 770 hours of audio and video data, covering various scenarios such as interviews, online/offline meetings, speeches, debates, daily conversations, and supports multiple languages including Chinese, English, Japanese and others. The dataset was constructed using automated methods, which combine audio and visual information to generate more accurate pseudo-labels, and conducts iterative training with a pre-trained speaker diarization model. The release of the M3SD dataset provides a new data resource for research on speaker diarization technology, and helps improve the generalization ability of models and their adaptability to different scenarios.
数据集概述
基本信息
- 许可证: openrail




