Speech Dataset for Neural Voice Watermark Removal
收藏资源简介:
Speech Dataset for Neural Voice Watermark Removal This record contains the clean, unwatermarked speech excerpts described in Section III.A ("Speech Database") of Exploiting Neural Speech Enhancement Models to Undermine Voice Watermarking (IEEE Transactions on Multimedia, 2026). The excerpts are intended for reproducible watermark embedding, watermark robustness, watermark-removal, and neural speech-enhancement experiments. The corpus is derived from the clean speech set of the Deep Noise Suppression (DNS) Challenge, originally sourced from LibriVox public-domain audiobooks. The release is a curated subset prepared for the associated watermarking study; it is not a redistribution of the full DNS corpus. Silero VAD was used to select utterances with at least 7 seconds of speech and internal pauses no longer than 500 ms, and each selected utterance was cropped to exactly 7.00 seconds. These strict validity criteria substantially constrain which parts of the source audiobooks can be used. The final selection balances the availability of valid excerpts and speaker representation: speakers with few valid candidates contribute only the excerpts that satisfy the criteria, while speakers with abundant valid material contribute at most four retained excerpts. Dataset contents and statistics The WAV files are distributed in three subset archives corresponding to training, validation, and test data: Training: 4842 WAV excerpts from 1256 speakers, 9.42 equivalent hours. Validation: 1203 WAV excerpts from 314 speakers, 2.34 equivalent hours. Test: 1202 WAV excerpts from 315 speakers, 2.34 equivalent hours. Total: 7247 WAV excerpts from 1885 speakers, 14.09 equivalent hours. All files are 16 kHz mono 16-bit PCM WAV files, each exactly 7.00 seconds long. The train, validation, and test subsets are speaker-disjoint. The release contains clean speech excerpts only; watermarked signals, embedded bitstreams, attack outputs, random payloads, transcripts, and trained model checkpoints are not included. Speaker metadata beyond the anonymous reader_id is not provided, and the speech domain is audiobook reading rather than conversational speech. Filename schema: book_<book_id>_chp_<chapter_id>_reader_<reader_id>_<sentence_id>-<excerpt_id>.wav. The DNS/LibriVox-derived fields identify the source book, chapter, reader, and sentence; excerpt_id identifies the retained 7-second excerpt generated in this release. The accompanying README_ZENODO.md provides the full dataset card, citation text, upstream references, and the per-speaker excerpt-distribution figure. The audio excerpts should be treated under the LibriVox/public-domain source terms documented by DNS. LibriVox states public-domain status under US law and advises users outside the US to verify local copyright status. Associated paper Please cite the accepted article associated with this dataset: A. M. Gomez, J. M. Martin-Donas, and A. M. Peinado, "Exploiting Neural Speech Enhancement Models to Undermine Voice Watermarking," IEEE Transactions on Multimedia, accepted for publication, 2026. Acknowledgments and funding The associated publication is part of the project PID2022-138711OB-I00 funded by MICIU/AEI/10.13039/501100011033 and by ERDF/EU. This dataset follows the Data Management Plan of the ASASVI project (Signal and Neural Processing against Spoofing Attacks and Deepfakes for Secure Voice Interaction, PID2022-138711OB-I00). Its ASASVI dataset identifier is Speech-Dataset-for-Neural-Voice-Watermark-Removal_250426_ASASVI_V1.0, and the corresponding DMP metadata record is included as Speech-Dataset-for-Neural-Voice-Watermark-Removal_250426_ASASVI_V1.0.xml.



