JusperLee/Hive-ALL
收藏资源简介:
Hive是一个用于数据高效查询式通用声音分离的高质量合成数据集。它包含2,442小时的原始音频和1,960万个混合音频,专门用于通用声音分离任务。数据集涵盖283个来自AudioSet本体的声音类别,采用语义一致的混合逻辑,采样率为44.1kHz。它通过自动管道构建,从无约束录音中挖掘高纯度单事件片段,并通过语义一致策略合成混合音频,以消除共现噪声。实验表明,尽管数据规模仅为百万小时基线的约0.2%,但在Hive上训练的模型实现了竞争性的分离准确性和感知质量,并在零样本泛化方面表现突出。数据集有两种发布格式:元数据版本(需本地重新生成混合音频)和预生成混合音频版本(可直接用于训练)。
Hive is a high-quality synthetic dataset for data-efficient query-based universal sound separation. It comprises 2,442 hours of raw audio and 19.6 million mixtures, designed for universal sound separation tasks. The dataset covers 283 sound categories from the AudioSet ontology, employs semantically consistent mixing logic, and has a sample rate of 44.1kHz. It is constructed via an automated pipeline that mines high-purity single-event segments from unconstrained recordings and synthesizes mixtures through semantically consistent strategies to eliminate co-occurrence noise. Experimental results demonstrate that, despite using only ~0.2% of the data scale of million-hour baselines, models trained on Hive achieve competitive separation accuracy and perceptual quality, with remarkable zero-shot generalization. The dataset is released in two formats: a metadata-only version (requiring local mixture regeneration) and a pre-generated mixed audio version (ready for training).



