radiotalk-us-audio-tada-noisy
收藏资源简介:
radiotalk-us-audio-tada-noisy 是一个专为航空通信自动语音识别(ASR)任务设计的带噪音频数据集。它是干净数据集 twangodev/radiotalk-us-audio-tada-clean 的 VHF AM 航空信道降级变体,旨在模拟真实航空通信环境中的信道退化效应。数据集包含 3,365,754 条带噪音频样本,由 1,121,918 条干净语音通过 3 次独立的信道模拟流程生成。每条样本通过一个概率性信道模拟管道处理,该管道根据 ATCO2 语料库的信噪比(SNR)分布(均值约 8 dB,范围 -5 至 +30 dB)进行校准,并遵循 ITU-R M.1084 / DO-186B 航空语音通带标准(300-3400 Hz,6阶巴特沃斯滤波器)。音频格式为 8 kHz 单声道 PCM_16 WAV,以字节形式嵌入在 Parquet 文件的 audio.bytes 字段中,数据被组织成 674 个约 470 MB 的分片。数据集模式是干净数据集的严格超集,保留了所有原始字段(如场景 ID、场景结构、说话人标签、原始和标准化文本、语音 ID、音频字节、令牌时间戳、上游模型信息等),并新增了与信道模拟相关的字段,包括干净行 ID、变体索引、随机种子、配置文件(飞行员或管制员)、应用的效果链、有效 SNR 以及管道版本和指纹。信道模拟流程包括一系列概率性应用的效果,如自动增益控制、饱和、硬剪辑、多径抖动(仅限飞行员)、粉红噪声(SNR 采样)、异频哨声、PTT 按键声、静噪尾音和编解码器往返等,所有处理均在 300-3400 Hz 频带内进行。该数据集适用于开发和评估在噪声和信道退化条件下(特别是航空 VHF AM 环境)的鲁棒 ASR 系统。
radiotalk-us-audio-tada-noisy is a noisy audio dataset specifically designed for automatic speech recognition (ASR) tasks in aviation communications. It is a VHF AM aviation channel degraded variant of the clean dataset twangodev/radiotalk-us-audio-tada-clean, aimed at simulating channel degradation effects in real-world aviation communication environments. The dataset contains 3,365,754 noisy audio samples, generated from 1,121,918 clean speech samples through three independent channel simulation processes. Each sample is processed via a probabilistic channel simulation pipeline calibrated based on the signal-to-noise ratio (SNR) distribution of the ATCO2 corpus (mean around 8 dB, range -5 to +30 dB) and adheres to the ITU-R M.1084 / DO-186B aviation voice passband standard (300-3400 Hz, 6th-order Butterworth filter). The audio format is 8 kHz mono PCM_16 WAV, embedded as bytes in the audio.bytes field of Parquet files, with data organized into 674 shards of approximately 470 MB each. The dataset schema is a strict superset of the clean dataset, preserving all original fields (such as scene ID, scene structure, speaker labels, original and normalized text, utterance ID, audio bytes, token timestamps, upstream model information, etc.) and adding new fields related to channel simulation, including clean row ID, variant index, random seed, profile (pilot or controller), applied effect chain, effective SNR, and pipeline version and fingerprint. The channel simulation process includes a series of probabilistically applied effects, such as automatic gain control, saturation, hard clipping, multipath jitter (pilot-only), pink noise (SNR sampling), heterodyne whistle, PTT keying, squelch tail, and codec round-trip, all processed within the 300-3400 Hz frequency band. This dataset is suitable for developing and evaluating robust ASR systems under noisy and channel-degraded conditions, particularly in aviation VHF AM environments.




