hoyyu1/infant-cry-detection
收藏资源简介:
本数据集专为嘈杂家庭环境下的婴儿哭声检测而设计,旨在提供一个多样化、具代表性且具备鲁棒性的音频样本集合,用于训练和评估机器学习模型,尤其是深度神经网络。数据集通过整合多个公开数据集构建,以确保样本的多样性和代表性。它包含约39.6小时的音频记录,划分为两类:婴儿哭声(正样本)和非哭声(负样本),并采用五折平衡划分供交叉验证。正样本数据来自CryCeleb2023、EnesBabyCries1和2020科大讯飞A.I.开发者大赛;负样本数据整合了VoxCeleb、ESC 50、Cat Meowing和DASEE数据集,并细分为语音、人类非语音、猫叫声、家庭噪音和静音五个子类。数据预处理包括移除静音片段、统一音频样本长度(5秒)和转换为单声道16 kHz WAV PCM格式。数据增强方法包括速度扰动、环境噪声破坏以及时间与频率掩蔽。数据集总计包含13,330个哭声样本和各类非哭声样本,具体数量为:语音4,963、家庭噪音4,506、人类非语音2,076、猫叫声1,953、静音1,650。
This dataset is designed for infant cry detection in noisy household environments. It aims to provide a diverse, representative, and robust collection of audio samples for training and evaluating machine learning models, particularly deep neural networks. The dataset was constructed by integrating multiple public datasets to ensure diversity and representativeness. It contains approximately 39.6 hours of audio recordings, partitioned into two classes: cry (positive) and non-cry (negative), with 5 balanced folds for cross-validation. The cry data subset consists of audio recordings from CryCeleb2023, EnesBabyCries1, and the 2020 iFLYTEK A.I. Developer Competition. The non-cry data integrates VoxCeleb, ESC 50, Cat Meowing, and DASEE datasets, and is divided into five subclasses: Speech, Human non-speech, Cat Meows, Household noise, and Silences. Data preprocessing includes removing silent segments, standardizing audio sample length to 5 seconds, and converting to single-channel 16 kHz WAV PCM format. Data augmentation methods include speed perturbation, environmental corruption, and time and frequency masking. The dataset consists of 13,330 cry samples and various non-cry samples, with specific counts: Speech 4,963, Household noise 4,506, Human non-speech 2,076, Cat Meows 1,953, and Silences 1,650.



