Dataset used in COALA: Co-Aligned Autoencoders for Learning Semantically Enriched Audio Representations
收藏资源简介:
This dataset consists of two hdf5 files that contain pre-computed log-mel spectrograms that have been used to to train audio embedding models. The dataset is split into a training set and a validation set containing respectively 170793 and 19103 spectrogram patches with their accompanying multi-hot encoded tags from a vocabulary of 1000 tags provided by Freesound users. More details can be found in "COALA: Co-Aligned Autoencoders for Learning Semantically Enriched Audio Representations" by X. Favory, K. Drossos, T. Virtanen, and X. Serra. The code is available at this GitHub repository. License: This dataset is derived from content from the Freesound collection. All sounds are released under Creative Commons (CC) licenses from either CC0, CC-BY, CC-S+, or CC-BY-NC. We attribute authors of all the sounds used in the dataset and provide their corresponding licenses in the attributions.txt file.
本数据集包含两个HDF5格式文件,其中存储了用于训练音频嵌入模型的预计算对数梅尔频谱图(log-mel spectrogram)。本数据集划分为训练集与验证集,分别包含170793和19103个频谱图块,附带由Freesound用户提供的1000个标签词汇表对应的多热编码标签。更多细节可参阅X. Favory、K. Drossos、T. Virtanen与X. Serra合著的论文《COALA:用于学习语义增强型音频表征的协同对齐自编码器》(COALA: Co-Aligned Autoencoders for Learning Semantically Enriched Audio Representations)。相关代码可在本GitHub仓库获取。许可声明:本数据集源自Freesound音频库的内容,所有音频素材均采用知识共享(Creative Commons,CC)许可协议发布,具体包括CC0、CC-BY、CC-S+及CC-BY-NC协议。本数据集已标注所用全部音频的原作者,并在attributions.txt文件中提供对应许可信息。



