遇见数据集

twangodev/radiotalk-us-audio-tada-clean

收藏
Hugging Face2026-05-20 更新2026-05-31 收录
官方服务:

资源简介:

该数据集名为RadioTalk US Audio (Clean),是一个合成清晰语音音频数据集,包含约10万个美国空中交通管制对话场景。每个场景以对话轮次为单位存储,音频为24 kHz单声道PCM_16 WAV格式。数据集总时长为2,180.4小时(约91天),包含99,957个场景(平均每个场景11.2轮对话,每轮平均7.0秒),使用了2,000个独特语音(通过Hume TADA零样本克隆生成)。数据来源于twangodev/radiotalk-us-transcripts-100k(转录文本)和twangodev/radiotalk-voices-2k(语音)。该数据集适用于文本转语音和自动语音识别任务,特别针对航空无线电通信领域。音频使用HumeAI/tada-3b-ml合成,基于Meta Llama 3.2 3B模型,遵循Llama 3.2社区许可协议。

The dataset is named RadioTalk US Audio (Clean), a synthesized clean-speech audio dataset for approximately 100k US air-traffic-control conversation scenarios. Each row represents one turn, with embedded 24 kHz mono PCM_16 WAV audio. It contains 2,180.4 hours of audio (~91 days), 99,957 scenarios (averaging 11.2 turns per scenario and 7.0 seconds per turn), and 2,000 unique voices (zero-shot cloned via Hume TADA). The data sources are transcripts from twangodev/radiotalk-us-transcripts-100k and voices from twangodev/radiotalk-voices-2k. It is designed for text-to-speech and automatic-speech-recognition tasks, specifically in aviation radio communication. The audio was synthesized using HumeAI/tada-3b-ml, built on Meta Llama 3.2 3B, and distributed under the Llama 3.2 Community License Agreement.

提供机构:
twangodev
二维码
社区交流群
二维码
科研交流群
商业服务