遇见数据集

BSD10k (Broad Sound Dataset 10k)

收藏
Zenodo2025-10-01 更新2026-05-26 收录
官方服务:

资源简介:

The BSD10k dataset (Broad Sound Dataset 10k) is an open collection of human-labeled sounds containing over 10k Freesound audio clips, annotated according to the 23 second-level classes defined in the Broad Sound Taxonomy. The dataset has been created at the Music Technology Group of Universitat Pompeu Fabra. Dataset characteristics The dataset consists of 10,309 sounds from Freesound, totaling ~32.5 hours of single-labeled audio. The sounds are cropped to a maximum length of 30 seconds, resulting in variable durations ranging from 0.01 to 30s. Audio lengths vary due to the heterogeneity of the sound classes and the range of contributions from Freesound users. The original files downloaded from Freesound are converted to a standardized format of uncompressed WAV files with 44.1 kHz sampling rate, 16-bit depth, and mono channel. The dataset’s audio files occupy approximately 7.9 GB and can be found in the audio folder. All sounds have been manually labeled by human annotators, with an estimated error rate of ~1%. The dataset categorizes the sounds into 23 classes, which are the second-level categories of the Broad Sound Taxonomy (see details below). The annotated data has a non-uniform distribution across these categories. For each audio file, the dataset includes the assigned category label, descriptive metadata (title, tags), provenance information (ID, uploader), and the license, all provided in BSD10k_metadata.csv. For more details on the dataset creation and its contents, please refer to our paper "Heterogeneous Sound Classification with the Broad Sound Taxonomy and Dataset", specifically Section 3.1. An overview of the BSD10k dataset is also available on the support site. Taxonomy The Broad Sound Taxonomy (BST) organizes sounds into a two-level hierarchical structure with 5 top-level and 23 second-level categories. The top-level categories cover distinct types of sounds: Music, Instrument samples, Speech, Sound effects, and Soundscapes. The taxonomy is designed to classify any type of sound while remaining broad, comprehensive, and easy to use. It can be used for organizing and filtering sounds in heterogeneous sound collections, such as Freesound, as well as in personal sound libraries. More details about the categories can be found in BST_description.csv, and additional information about the taxonomy is provided in our upcoming journal paper "A General-Purpose Sound Taxonomy for the Classification of Heterogeneous Sound Collections". Citation When using all or part of the BSD10k dataset, please cite our paper (available from [UPF e-repositori] [arXiv] [DCASE2024 proceedings]): @inproceedings{anastasopoulou2024heterogeneous, title = {Heterogeneous Sound Classification with the {{Broad Sound Taxonomy}} and {{Dataset}}}, author = {Anastasopoulou, Panagiota and Torrey, Jessica and Serra, Xavier and Font, Frederic}, booktitle = {Workshop on {{Detection}} and {{Classification}} of {{Acoustic Scenes}} and {{Events}} ({{DCASE}})}, year = {2024}} License BSD10k is released in its entirety under the CC BY 4.0 license. We note, though, that each audio file is released under its own Creative Commons (CC) license, as defined by the respective uploader in Freesound. Some sounds require attribution to their original authors, while others forbid commercial reuse. If the dataset is used in a commercial setting, the sounds with CC BY-NC licenses should be excluded. This is the distribution of sounds per license: CC0: 3,187 CC BY: 5,534 CC BY-NC: 1,192 CC Sampling+: 396 Links to the license deeds for each sound can be further accessed through BSD10k_metadata.csv. Data structure BSD10k can be accessed as follows: 𝚛𝚘𝚘𝚝/├── 𝚊𝚞𝚍𝚒𝚘/ 𝙰𝚞𝚍𝚒𝚘 𝚏𝚒𝚕𝚎𝚜├── 𝚖𝚎𝚝𝚊𝚍𝚊𝚝𝚊/ 𝙼𝚎𝚝𝚊𝚍𝚊𝚝𝚊 𝚏𝚒𝚕𝚎𝚜│ ├── 𝙱𝚂𝙳𝟷0𝚔_𝚖𝚎𝚝𝚊𝚍𝚊𝚝𝚊.𝚌𝚜𝚟 𝙳𝚊𝚝𝚊𝚜𝚎𝚝'𝚜 𝚖𝚎𝚝𝚊𝚍𝚊𝚝𝚊│ ├── 𝙱𝚂𝚃_𝚍𝚎𝚜𝚌𝚛𝚒𝚙𝚝𝚒𝚘𝚗.𝚌𝚜𝚟 𝚃𝚊𝚡𝚘𝚗𝚘𝚖𝚢 𝚒𝚗𝚏𝚘𝚛𝚖𝚊𝚝𝚒𝚘𝚗│ └── 𝙱𝚂𝚃_𝚍𝚒𝚊𝚐𝚛𝚊𝚖.𝚙𝚗𝚐 𝚃𝚊𝚡𝚘𝚗𝚘𝚖𝚢 𝚍𝚒𝚊𝚐𝚛𝚊𝚖└── 𝚁𝙴𝙰𝙳𝙼𝙴.𝚖𝚍 𝙳𝚘𝚌𝚞𝚖𝚎𝚗𝚝𝚊𝚝𝚒𝚘𝚗 (𝚝𝚑𝚊𝚝 𝚢𝚘𝚞 𝚊𝚛𝚎 𝚗𝚘𝚠 𝚛𝚎𝚊𝚍𝚒𝚗𝚐) BSD10k_metadata.csv is the main metadata file, containing annotations and additional information for each sound. Each row corresponds to one sound and includes the following fields: sound_id: Freesound ID used as the unique identifier of the sound. The audio files found in the audio folder are named using this ID, with a .wav extension for the audio format. class: Second-level class code of the sound. class_idx: Second-level class index (0-22), ordered according to the taxonomy. class_top: Corresponding top-level class code. uploader: User who uploaded the sound in Freesound. license: Link to the license of the sound. title: Sound title provided by the uploader. tags: Tags associated with the sound provided by the uploader. The mapping of class codes to their corresponding full class names can be found in BST_description.csv, which also includes a description and examples for each class. A diagram of the taxonomy (BST_diagram.png) is also included for a quick overview of the categories. Acknowledgments This research is partially funded by the Generalitat de Catalunya (2023FI-100252, Joan Oró program) and the IA y Música Cátedra (TSI-100929-2023-1, Cátedras ENIA 2022, SE Digitalización e IA, EU NGEU). Contact You are welcome to contact Panagiota Anastasopoulou if you have any questions, at panagiota.anastasopoulou@upf.edu.

BSD10k数据集(Broad Sound Dataset 10k,宽声数据集10k)是一个由人工标注的开源音频集合,包含超过1万条Freesound平台音频片段,其标注依据宽声分类体系(Broad Sound Taxonomy, BST)所定义的23个二级类别。该数据集由庞培法布拉大学(Universitat Pompeu Fabra, UPF)音乐技术组研发构建。 ### 数据集特性 本数据集包含来自Freesound平台的10309条音频,总时长约32.5小时,均为单标签标注音频。所有音频均被裁剪至最长30秒,因此实际时长介于0.01秒至30秒之间。音频时长的差异源于音频类别的异质性以及Freesound平台用户上传内容的多样性。从Freesound平台下载的原始音频文件均被转换为标准化格式:未压缩WAV格式,采样率44.1kHz,位深16比特,单声道。数据集的音频文件总占用空间约7.9GB,存储于audio文件夹中。 所有音频均由人工标注员手动标注,标注错误率约为1%。数据集将音频划分为23个类别,即宽声分类体系的二级类别(详见下文)。标注数据在各类别间分布不均。每条音频文件均包含所属类别标签、描述性元数据(标题、标签)、来源信息(ID、上传者)以及授权协议,所有信息均存储于BSD10k_metadata.csv文件中。如需了解数据集构建与内容的更多细节,请参阅我们的论文"Heterogeneous Sound Classification with the Broad Sound Taxonomy and Dataset"的3.1章节。BSD10k数据集的概览也可在支持站点获取。 ### 分类体系 宽声分类体系(Broad Sound Taxonomy, BST)采用两级层级结构组织音频,包含5个一级类别与23个二级类别。一级类别涵盖了所有类型的音频:音乐(Music)、乐器采样(Instrument samples)、语音(Speech)、音效(Sound effects)以及声景(Soundscapes)。该分类体系旨在实现全类型音频的分类,同时保持覆盖范围广泛、内容全面且易于使用。它可用于对异质音频集合(如Freesound平台)以及个人音频库中的音频进行组织与筛选。各类别的详细信息可参阅BST_description.csv文件,有关该分类体系的更多内容将发表于我们即将刊出的期刊论文"A General-Purpose Sound Taxonomy for the Classification of Heterogeneous Sound Collections"。 ### 引用方式 当使用BSD10k数据集的全部或部分内容时,请引用我们的论文(可从[UPF电子仓储库]、[arXiv]、[DCASE2024会议论文集]获取): bibtex @inproceedings{anastasopoulou2024heterogeneous, title = {Heterogeneous Sound Classification with the {{Broad Sound Taxonomy}} and {{Dataset}}}, author = {Anastasopoulou, Panagiota and Torrey, Jessica and Serra, Xavier and Font, Frederic}, booktitle = {Workshop on {{Detection}} and {{Classification}} of {{Acoustic Scenes}} and {{Events}} ({{DCASE}})}, year = {2024} } ### 授权协议 BSD10k数据集整体采用CC BY 4.0协议发布。但需注意,每条音频文件的实际授权协议由其在Freesound平台的上传者指定,采用各自的知识共享(Creative Commons, CC)协议。部分音频需标注原作者,部分则禁止商业复用。若将数据集用于商业场景,需排除采用CC BY-NC协议的音频文件。 各授权协议对应的音频数量分布如下: - CC0:3187条 - CC BY:5534条 - CC BY-NC:1192条 - CC Sampling+:396条 每条音频的授权协议链接可通过BSD10k_metadata.csv文件获取。 ### 数据结构 BSD10k数据集的目录结构如下: root/ ├── audio/ # 音频文件 ├── metadata/ # 元数据文件 │ ├── BSD10k_metadata.csv # 数据集元数据 │ ├── BST_description.csv # 分类体系信息 │ └── BST_diagram.png # 分类体系示意图 └── README.md # 说明文档(即当前阅读的内容) BSD10k_metadata.csv是核心元数据文件,存储了每条音频的标注与附加信息。文件中每一行对应一条音频,包含以下字段: - sound_id:音频的唯一标识符,即Freesound平台的音频ID。audio文件夹中的音频文件均以此ID命名,扩展名为.wav。 - class:音频所属的二级类别代码。 - class_idx:二级类别索引(0-22),按照分类体系的顺序排序。 - class_top:对应的一级类别代码。 - uploader:该音频在Freesound平台的上传者用户名。 - license:该音频的授权协议链接。 - title:上传者提供的音频标题。 - tags:上传者为该音频添加的关联标签。 类别代码与完整类别名称的映射关系可参阅BST_description.csv文件,该文件同时包含各类别的描述与示例。此外,数据集还提供了分类体系示意图(BST_diagram.png),可快速概览所有类别。 ### 致谢 本研究受加泰罗尼亚政府(项目编号2023FI-100252,Joan Oró计划)以及IA y Música教席项目(项目编号TSI-100929-2023-1,Cátedras ENIA 2022,西班牙数字化与人工智能秘书处,欧盟下一代欧盟计划)资助。 ### 联系方式 如有任何疑问,可联系Panagiota Anastasopoulou,邮箱:panagiota.anastasopoulou@upf.edu。

提供机构:
Zenodo
创建时间:
2025-10-01
二维码
社区交流群
二维码
科研交流群
商业服务