AVASpeech-SMAD
收藏资源简介:
AVASpeech-SMAD数据集由佐治亚理工学院音乐技术中心创建,旨在支持语音和音乐活动检测(SMAD)研究。该数据集是对原有AVASpeech数据集的扩展,增加了帧级别的音乐标签,使得数据集成为首个包含音乐和语音强多音标签的开源数据集。数据集包含160个15分钟的YouTube视频片段,总时长45小时,涵盖多种内容、语言、流派和制作质量。数据集的创建过程包括手动标注和验证,通过迭代交叉检查和简单的自动检查来确保标签质量。该数据集适用于训练和评估未来的SMAD系统,特别是在解决现实世界中语音和音乐共存问题方面具有重要意义。
The AVASpeech-SMAD dataset was developed by the Georgia Tech Center for Music Technology to support research on Speech and Music Activity Detection (SMAD). As an extended version of the original AVASpeech dataset, it adds frame-level music annotations, making it the first open-source dataset featuring strong polyphonic labels for both speech and music. The dataset consists of 160 15-minute YouTube video clips, with a total duration of 45 hours, covering a wide range of content, languages, music genres, and production qualities. Its creation process involves manual annotation and validation, utilizing iterative cross-checks and basic automatic checks to ensure the quality of the labels. This dataset can be used to train and evaluate future SMAD systems, and is particularly valuable for addressing the real-world challenge of coexisting speech and music.



