遇见数据集

DnR-nonverbal-v2 dataset

收藏
Zenodo2025-09-09 更新2026-05-26 收录
官方服务:

资源简介:

Introduction DnR-nonverbal is a dataset for cinematic audio source separation (CASS) based on Divide and Remaster (DnR) dataset. Unlike conventional datasets, our dataset contains non-verbal sounds such as laughter and screaming, just like actual movie audio. Our dataset enables CASS models to allocate non-verbal sounds to the same stem as speech. Examples of clips and separation results are available at https://tky823.github.io/hasumi2025dnr.github.io/ for DnR-nonverbal-v1. Update DnR-nonverbal-v2 is updated from DnR-nonverbal-v1 in the following points: Based on the DnR dataset, the audio quantization bit depth has been set to 32 bits. By majority vote among the three annotators, samples not originating from human voices were removed. New samples crawled from FreeSound. These samples are not necessarily CC0. Please refer to licenses.tsv. New "sigh" samples from VocalSound dataset. How to Use Download dnr-nonverbal.tar.gz from this page. Extract dnr-nonverbal.tar.gz by tar xvzf dnr-nonverval-v2.tar.gz (optional) Mix directories with the DnR. Our sample IDs are assigned in such a way that they do not duplicate DnR. Dataset Structure The dataset structure is based on DnR, except that our dataset contains non-verbal sounds as a part of the speech stem. dnr-nonverbal-v2 ├── tr │ ├── 100009 │ │ ├── annots.csv │ │ ├── background.wav │ │ ├── foreground.wav │ │ ├── mix.wav │ │ ├── music.wav │ │ ├── nonverbal.wav │ │ ├── reading.wav │ │ ├── sfx.wav │ │ └── speech.wav │ ├── 100031 │ ... ├── cv └── tt reading.wav: Reading style speech extracted from LibriSpeech. nonverbal.wav: Non-verbal sounds collected from FSD50K and newly crawled from FreeSound. In addition, sign samples are taken from VocalSound. speech.wav: Mixture of reading style speech and non-verbal sounds. music.wav: Background music extracted from FMA (medium). foreground.wav: Foreground effect sounds collected from FSD50K. background.wav: Background effect sounds collected from FSD50K. sfx.wav: Foreground and background effect sounds. annots.csv: A metadata file that identifies sources of sounds. Citation

提供机构:
Zenodo
创建时间:
2025-09-09
二维码
社区交流群
二维码
科研交流群
商业服务