TTS-AGI/advanced-soundscapes-stage-1
收藏资源简介:
该数据集是LAION通用音频标注管道(UAAP)数据生成计划的第一阶段输出,名为“高级声景第一阶段——原始组件(5M)”。它包含5000个分片,总计500万个声景配方,每个配方包含原始音频组件:语音(spkN.flac/json)、音乐(musicN.flac/json)、音效(sfxN.flac/json)和人声爆发(vbN.flac/json),以及完整的配方元数据(recipe.json)。数据格式为WebDataset tar分片,每个分片包含1000个片段。源数据集包括laion/majestrino-data(语音,约50种语言,CC0许可证)、laion/captioned-ai-music-snippets(音乐,3-30秒,Apache-2.0许可证)、laion/generated-sound-events(音效,合成,1190个类别,Apache 2.0许可证)和laion/synthetic_vocal_bursts(人声爆发,宽松许可证)。第一阶段流水线在CPU上运行,包括采样配方、从源数据集抽取原始音频、解码并重采样为16 kHz单声道FLAC、计算精确混合计划(包括时间、响度、说话者ID)、收集完整源元数据,并写入WebDataset行。数据集还定义了多种配方原型(如平衡三重奏、语音主导、音乐主导等),以及响度梯度和持续时间频段,用于生成多样化的声景。
This dataset contains Stage 1 output from the LAION Universal Audio Annotation Pipeline (UAAP) data generation plan, titled Advanced Soundscapes Stage 1 — Raw Components (5M). It includes 5,000 shards with 5,000,000 soundscape recipes, each containing raw audio components: speech (spkN.flac/json), music (musicN.flac/json), sound effects (sfxN.flac/json), and vocal bursts (vbN.flac/json), along with full recipe metadata (recipe.json). The format is WebDataset .tar shards, with 1000 clips per shard. Source datasets are laion/majestrino-data (speech, ~50 languages, CC0 license), laion/captioned-ai-music-snippets (music, 3–30s, Apache-2.0 license), laion/generated-sound-events (SFX, synthetic, 1190 classes, Apache 2.0 license), and laion/synthetic_vocal_bursts (vocal bursts, permissive license). The Stage 1 pipeline is CPU-only and involves sampling a recipe, drawing raw audio pieces from source datasets, decoding and resampling to 16 kHz mono FLAC, computing an exact mix plan (ground-truth timings, loudness, speaker IDs), gathering full source metadata, and writing WebDataset rows. The dataset also defines recipe archetypes (e.g., balanced trio, speech-dominant, music-dominant), a loudness ladder, and duration bands for generating diverse soundscapes.




