遇见数据集

EnvSDD/EnvSDD

收藏
Hugging Face2026-05-17 更新2026-05-31 收录
官方服务:

资源简介:

EnvSDD是首个用于环境声音深度伪造检测的大规模精选数据集。随着音频生成模型能够产生极其逼真的声景,这为错误信息和公共安全带来了新的风险。尽管语音和歌唱深度伪造检测已受到广泛研究关注,但环境声音深度伪造仍是一个未解决的问题。EnvSDD填补了这一空白,包含45.25小时的真实环境音频和316.74小时的AI生成(深度伪造)音频,涵盖单声道和多声道条件,使用5个文本到音频(TTA)模型和2个音频到音频(ATA)模型生成,并包含4种测试条件(域内和具有挑战性的域外场景)。所有音频为16 kHz、4秒/片段。数据集包含训练集(139,055个样本)、验证集(39,710个样本)、测试集(39,768个样本)和剩余集(107,259个样本),总计超过30万样本。数据字段包括音频、标签(真实/伪造)、攻击类型(TTA/ATA/-)、生成模型(如AudioLDM、TangoFlux)、源数据集(如UrbanSound8K)、文件名、场景标签(多声道片段)、事件标签(单声道片段)和字幕(用于TTA生成的文本描述)。数据集基于多个真实音频源(UrbanSound8K、DCASE 2023、TAU UAS 2019等)构建,并支持环境声音深度伪造检测的研究和评估。

EnvSDD is the first large-scale curated dataset for environmental sound deepfake detection. As audio generation models can produce highly realistic soundscapes, this poses new risks to misinformation and public safety. While speech and singing deepfake detection has garnered extensive research attention, environmental sound deepfake detection remains an unsolved challenge. EnvSDD fills this gap, containing 45.25 hours of real environmental audio and 316.74 hours of AI-generated (deepfake) audio across mono and multi-channel conditions. It is generated using 5 text-to-audio (TTA) models and 2 audio-to-audio (ATA) models, and includes 4 test scenarios (in-domain and challenging out-of-domain settings). All audio clips are sampled at 16 kHz with a duration of 4 seconds each. The dataset is split into a training set (139,055 samples), validation set (39,710 samples), test set (39,768 samples), and a residual set (107,259 samples), totaling over 300,000 samples. The data fields include audio, label (real/fake), attack type (TTA/ATA/-), generation model (e.g., AudioLDM, TangoFlux), source dataset (e.g., UrbanSound8K), file name, scene label (for multi-channel clips), event label (for mono-channel clips), and caption (text description used for TTA generation). Built upon multiple real audio sources including UrbanSound8K, DCASE 2023, TAU UAS 2019, and others, this dataset supports research and evaluation for environmental sound deepfake detection.

提供机构:
EnvSDD
二维码
社区交流群
二维码
科研交流群
商业服务