遇见数据集

GARAGEM: General Automotive Real and Artificial speech corpus for Garage Environments and Maintenance in brazilian portuguese

收藏
Zenodo2025-12-16 更新2026-05-26 收录
官方服务:

资源简介:

General Description This dataset is a collection of audio recordings focused on the automotive workshop domain, consisting of a real dataset extracted from YouTube and four synthetic datasets generated by different text-to-speech APIs. The total dataset contains approximately 60 hours of audio. Dataset Structure The dataset is organized into five main subsets: 1. YouTube Dataset (Real) Contains 4,412 audio segments (9 hours and 45 minutes) extracted from YouTube channels focused on automotive workshop topics. Split: - Training: 3,121 segments (6 hours and 51 minutes) - Test: 1,291 segments (2 hours and 54 minutes) Metadata Files: - train_metadata.csv - test_metadata.csv Metadata columns: - url: YouTube video URL - start_timestamp: Segment start time - end_timestamp: Segment end time - human_transcription: Human-annotated transcription 2. Google Cloud Dataset (Synthetic) Contains 6,367 synthetic audio segments generated by Google Cloud API. Format: WAV, 16 KHz Available Metadata Files: - metadata_100.csv: 100% of audios (11 hours and 15 minutes) - metadata_50.csv: 50% of audios (5 hours and 36 minutes) - metadata_20.csv: 20% of audios (2 hours and 14 minutes) - metadata_10.csv: 10% of audios (1 hour and 7 minutes) - metadata_5.csv: 5% of audios (33 minutes) Metadata columns: - audio_filename: Audio file - transcription: Transcription used to generate the audio - speaker: Speaker voice defined in generation configuration 3. Gemini Dataset (Synthetic) Contains 6,367 synthetic audio segments generated by Google API with Gemini assistance. Format: WAV, 16 KHz Available Metadata Files: - metadata_100.csv: 100% of audios (12 hours and 49 minutes) - metadata_50.csv: 50% of audios (6 hours and 24 minutes) - metadata_20.csv: 20% of audios (2 hours and 33 minutes) - metadata_10.csv: 10% of audios (1 hour and 16 minutes) - metadata_5.csv: 5% of audios (37 minutes) Metadata columns: - audio_filename: Audio file - transcription: Transcription used to generate the audio - speaker: Speaker voice defined in generation configuration 4. Google Cloud with Noise Dataset (Synthetic) Contains 6,367 synthetic audio segments generated by Google Cloud API with added background noise. Format: WAV, 16 KHz Available Metadata Files: - metadata_100.csv: 100% of audios (11 hours and 15 minutes) - metadata_50.csv: 50% of audios (5 hours and 36 minutes) - metadata_20.csv: 20% of audios (2 hours and 14 minutes) - metadata_10.csv: 10% of audios (1 hour and 7 minutes) - metadata_5.csv: 5% of audios (33 minutes) Metadata columns: - audio_filename: Audio file - transcription: Transcription used to generate the audio - speaker: Speaker voice defined in generation configuration - audio_with_noise_from_musan: Noise file used for noise addition - noise_window: Time window extracted from the noise audio - noise_scale: Scale of added noise Noise addition details: Noise was added using the MUSAN dataset (https://www.openslr.org/17/), which contains various types of noise. The noise files, time windows, and scales were randomly chosen for each audio. 5. Gemini with Noise Dataset (Synthetic) Contains 6,367 synthetic audio segments generated by Google API with Gemini assistance, with added background noise. Format: WAV, 16 KHz Available Metadata Files: - metadata_100.csv: 100% of audios (12 hours and 49 minutes) - metadata_50.csv: 50% of audios (6 hours and 24 minutes) - metadata_20.csv: 20% of audios (2 hours and 33 minutes) - metadata_10.csv: 10% of audios (1 hour and 16 minutes) - metadata_5.csv: 5% of audios (37 minutes) Metadata columns: - audio_filename: Audio file - transcription: Transcription used to generate the audio - speaker: Speaker voice defined in generation configuration - audio_with_noise_from_musan: Noise file used for noise addition - noise_window: Time window extracted from the noise audio - noise_scale: Scale of added noise Noise addition details: Noise was added using the MUSAN dataset (https://www.openslr.org/17/), which contains various types of noise. The noise files, time windows, and scales were randomly chosen for each audio. Automotive Domain Dictionary The dataset includes a dictionary file (`dictionary.csv`) containing automotive workshop terminology used for generating synthetic sentences. File: dictionary.csv Columns: - term: Automotive workshop term - meaning: Definition of the term (when necessary to assist in generation) - source: Term category Term Distribution: - Automotive parts: 330 terms - Generic term: 142 terms - Car model: 136 terms - Automotive brands: 113 terms Total terms: 721 This dictionary was used as the basis for generating the synthetic sentences present in the dataset, ensuring domain-specific vocabulary coverage. File Organization garagem/ ├── dictionary.csv ├── youtube/ │ ├── train_metadata.csv │ └── test_metadata.csv ├── google_cloud/ │ ├── audios/ │ ├── metadata_100.csv │ ├── metadata_50.csv │ ├── metadata_20.csv │ ├── metadata_10.csv │ └── metadata_5.csv ├── gemini/ │ ├── audios/ │ ├── metadata_100.csv │ ├── metadata_50.csv │ ├── metadata_20.csv │ ├── metadata_10.csv │ └── metadata_5.csv ├── google_cloud_with_noise/ │ ├── audios/ │ ├── metadata_100.csv │ ├── metadata_50.csv │ ├── metadata_20.csv │ ├── metadata_10.csv │ └── metadata_5.csv └── gemini_with_noise/ ├── audios/ ├── metadata_100.csv ├── metadata_50.csv ├── metadata_20.csv ├── metadata_10.csv └── metadata_5.csv Technical Specifications Audio Format (Synthetic Datasets): - Format: WAV - Sample Rate: 16 KHz Metadata File Format: - Format: CSV - Encoding: UTF-8 Recommended Use This dataset is suitable for: - Training automatic speech recognition (ASR) models - Research in speech processing in the automotive workshop domain - Evaluating model robustness to noise - Comparing real and synthetic data - Domain-specific vocabulary analysis Citation If you use this dataset in your research, please cite appropriately. License Creative Commons Attribution 4.0 International

提供机构:
Zenodo
创建时间:
2025-12-16
二维码
社区交流群
二维码
科研交流群
商业服务