VoxEffects
收藏资源简介:
VoxEffects是由日本国立情报学研究所开发的语音音频效果数据集,旨在支持语音音频效果识别研究。该数据集基于干净语音录音构建,包含2520种预设组合,覆盖六种常见音频效果(如降噪、动态范围压缩等),并提供了多粒度监督信息。数据集通过可扩展的渲染管道生成,支持离线合成和实时渲染,适用于训练和评估。VoxEffects主要应用于语音处理领域,旨在解决音频效果识别问题,包括效果存在检测、预设分类和强度预测等任务,为语音内容理解、音频工程辅助和音频取证等应用提供支持。
VoxEffects is a speech audio effect dataset developed by the National Institute of Informatics (NII) of Japan, designed to support research on speech audio effect recognition. Constructed from clean speech recordings, this dataset contains 2520 preset combinations, covers six common audio effects such as noise reduction, dynamic range compression, etc., and provides multi-granularity supervision information. The dataset is generated via a scalable rendering pipeline, supports both offline synthesis and real-time rendering, and is suitable for model training and evaluation. VoxEffects is primarily applied in the field of speech processing, aiming to address audio effect recognition tasks including effect presence detection, preset classification, and intensity prediction, providing support for applications such as speech content understanding, audio engineering assistance, and audio forensics.




