AudioSetMix
收藏资源简介:
AudioSetMix是一个高质量的音频-文本数据集,由普林斯顿大学创建。该数据集通过将音频变换应用于AudioSet的剪辑,并结合大型语言模型(LLM)生成的自然语言描述,形成了音频与文本的配对。数据集包含49,971对音频-文本数据,支持多种音频变换,如速度、音高、音量和持续时间,以及混合和串联变换。这些变换使得数据集能够支持文本引导的音频编辑研究,并提供原始和编辑后的音频数据。AudioSetMix数据集的应用领域包括文本到音频的检索和模型对音频事件修饰符的理解,旨在解决现有音频-语言数据集中修饰符(如形容词和副词)缺失的问题。
AudioSetMix is a high-quality audio-text dataset created by Princeton University. This dataset generates audio-text pairs by applying audio transformations to AudioSet clips and combining them with natural language descriptions generated by Large Language Models (LLMs). It contains 49,971 audio-text pairs, supporting a variety of audio transformations including speed, pitch, volume, duration, mixing and concatenation operations. These transformations enable the dataset to support text-guided audio editing research, providing both original and edited audio data. The application fields of AudioSetMix include text-to-audio retrieval and model comprehension of audio event modifiers, aiming to solve the problem of missing modifiers such as adjectives and adverbs in existing audio-language datasets.



