Labeled Voice Audio Dataset
收藏资源简介:
Labeled Voice Audio Dataset Overview This dataset contains a comprehensive collection of audio files and associated metadata for scientific research purposes. The labels include the voice quality assessments on the grbas scale. The dataset is structured to support various audio analysis, machine learning, and acoustic research applications. Dataset Contents Files Included - database.sql - Complete database dump containing all metadata, user information, and file references - data.zip - Compressed archive containing all audio files (1,374 audio files in WAV format) - README.md - Readme file with further information Database Structure The database contains the following key tables: - patients - Patient demographic information including gender, age and microphone type - sound_files - Audio file metadata including filenames, sound types and patient associations - users - User account information for system access and labeling activities - labelings - Voice quality assessments using the GRBAS scale (Grade, Roughness, Breathiness, Asthenia, Strain) - labeling_notes - Additional notes and comments associated with voice quality assessments Audio Data The audio collection consists of 1,374 individual audio files in WAV format. All files are contained within the 'data.zip' archive and can be extracted to access the individual audio samples. Audio Specifications: - Format: WAV (Waveform Audio File Format) - Total Files: 1,374 Data Sources: This dataset includes audio files from multiple sources: - Original recordings and assessments - Selected audio files from the Saarbrücken Voice Database (see Attribution section below) Citation If you use this dataset in your research, please cite it appropriately: VReedback GmbH & Co. KG (2025). Audio Research Dataset. Zenodo. https://doi.org/10.5281/zenodo.17080420 Additionally, please cite the Saarbrücken Voice Database: Pützer, M., & Koreman, J. (1997). A German Database of Patterns of Pathological Vocal Fold Vibration. In Phonus Research Report (Nummer 3, S. 143–153). Zenodo. https://doi.org/10.5281/zenodo.16258835 License This dataset is licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0). Attribution Saarbrücken Voice Database Portions of the audio data in this dataset are derived from the Saarbrücken Voice Database, which is also licensed under CC BY 4.0. When using this dataset, please also cite the original source: Pützer, M., & Koreman, J. (1997). A German Database of Patterns of Pathological Vocal Fold Vibration. In Phonus Research Report (Nummer 3, S. 143–153). Zenodo. https://doi.org/10.5281/zenodo.16258835



