AudioSetCaps|音频-语言多模态数据集|多模态数据数据集
收藏AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models
数据集概述
- 数据来源: AudioSet, YouTube-8M, VGGSound
- 音频文件数量: 6,117,099个10秒音频文件
- Q&A数据对数量: 18,414,789对
数据集内容
- 音频描述: 每个音频文件附有一个描述性标题。
- Q&A对: 每个音频文件附有三个Q&A对,作为生成最终标题的元数据。
示例
ID | 音频 | 描述 | Q&A 1 | Q&A 2 | Q&A 3 |
---|---|---|---|---|---|
_7Xe9vD3Hpg_4_10 | <audio controls><source src="Example /_7Xe9vD3Hpg_4_10.mp3" type="audio/mpeg"> Your browser does not support the audio element.</audio> | A solemn instrumental piece unfolds, featuring the melancholic strains of a cello and the resonant tolling of a bell. The initial tempo is slow and deliberate, gradually building intensity with each successive bell ring. | Question: Describe this audio according to the sounds in it. Answer: The music starts with a slow string melody and continues with a bass note. The sound of a bell rings and the music becomes more intense. | Question: Based on the QAs, give some information about the speech, such as the emotion of the speaker, the gender of the speaker, and the spoken language, only if speech is present in this audio. Answer: Im sorry, but there is no speech in the audio. | Question: Based on the QAs, give some information about the music, such as music genre and music instruments, only if music is present in this audio. Answer: The music genre is instrumental. The music instruments are the cello and the bell. |
-TL8Mp3xcUM_0_10 | <audio controls><source src="Example/-TL8Mp3xcUM_0_10.mp3" type="audio/mpeg"> Your browser does not support the audio element.</audio> | A woman expresses strong emotions with a noticeably high-pitched vocal tone. | Question: Describe this audio according to the sounds in it. Answer: A woman speaks with a high-pitched voice. | Question: Based on the QAs, give some information about the speech, such as the emotion of the speaker, the gender of the speaker, and the spoken language, only if speech is present in this audio. Answer: The speech is emotional, as the woman speaks in a high-pitched voice. | Question: Based on the QAs, give some information about the music, such as music genre and music instruments, only if music is present in this audio. Answer: There is no music in this audio. |
数据统计
数据集 | 音频描述数量 | Q&A描述数量 | 总计 |
---|---|---|---|
AudioSetCaps | 1910920 | 5736072 | 7646992 |
YouTube-8M | 4023990 | 12086037 | 16110027 |
VGGSound | 182189 | 592680 | 774869 |
总计 | 6117099 | 18414789 | 24531888 |
下载
许可证
- 仅允许学术用途
引用
bibtex @inproceedings{ bai2024audiosetcaps, title={AudioSetCaps: Enriched Audio Captioning Dataset Generation Using Large Audio Language Models}, author={JISHENG BAI and Haohe Liu and Mou Wang and Dongyuan Shi and Wenwu Wang and Mark D Plumbley and Woon-Seng Gan and Jianfeng Chen}, booktitle={Audio Imagination: NeurIPS 2024 Workshop AI-Driven Speech, Music, and Sound Generation}, year={2024}, url={https://openreview.net/forum?id=uez4PMZwzP} }

中国气象数据
本数据集包含了中国2023年1月至11月的气象数据,包括日照时间、降雨量、温度、风速等关键数据。通过这些数据,可以深入了解气象现象对不同地区的影响,并通过可视化工具揭示中国的气温分布、降水情况、风速趋势等。
github 收录
FAOSTAT Forestry
FAOSTAT Forestry数据集包含了全球森林资源的相关统计数据,涵盖了森林面积、木材产量、森林管理等多个方面。该数据集提供了详细的国别数据,帮助用户了解全球森林资源的现状和变化趋势。
www.fao.org 收录
Breast Cancer Dataset
该项目专注于清理和转换一个乳腺癌数据集,该数据集最初由卢布尔雅那大学医学中心肿瘤研究所获得。目标是通过应用各种数据转换技术(如分类、编码和二值化)来创建一个可以由数据科学团队用于未来分析的精炼数据集。
github 收录
NREL Wind Integration National Dataset (WIND) Toolkit
NREL Wind Integration National Dataset (WIND) Toolkit 是一个包含美国大陆风能资源和电力系统集成数据的综合数据集。该数据集提供了高分辨率的风速、风向、风能密度、电力输出等数据,覆盖了美国大陆的多个地理区域。这些数据有助于研究人员和工程师进行风能资源评估、电力系统规划和集成研究。
www.nrel.gov 收录
NASA Battery Dataset
用于预测电池健康状态的数据集,由NASA提供。
github 收录