PodcastFillers
收藏资源简介:
PodcastFillers是由罗切斯特大学和Adobe研究院合作创建的大型语音数据集,专注于填充词(如‘uh’或‘um’)的检测与分类。该数据集包含来自199个公共播客节目的145小时语音数据,涵盖超过350名不同性别和背景的说话者。数据集通过结合语音活动检测(VAD)模型和自动语音识别(ASR)系统,自动生成填充词候选,并通过众包方式进行手动验证和标注,总计包含35000个填充词标注和50000个其他语音事件标注。PodcastFillers旨在为填充词自动检测和分类提供一个基准数据集,以加速媒体编辑中的填充词处理任务,提高语音内容创作的效率。
PodcastFillers is a large-scale speech dataset co-created by the University of Rochester and Adobe Research, focusing on the detection and classification of filled pauses (e.g., 'uh' or 'um'). This dataset includes 145 hours of speech data from 199 public podcast episodes, involving over 350 speakers with diverse genders and backgrounds. It automatically generates filled pause candidates by combining Voice Activity Detection (VAD) models and Automatic Speech Recognition (ASR) systems, and then conducts manual verification and annotation via crowdsourcing. In total, it contains 35,000 filled pause annotations and 50,000 annotations for other speech events. PodcastFillers aims to provide a benchmark dataset for automatic filled pause detection and classification, so as to accelerate the processing of filled pauses in media editing and improve the efficiency of speech content creation.




