Quadruplet Dataset
收藏资源简介:
该数据集由加州大学伯克利分校的研究团队构建,主要用于音频纹理操作任务。数据集结合了来自LibriSpeech和VCTK的语音数据以及BBC SFX的环境音效,每个样本包含一个示例输入、示例输出、新输入音频和转换后的输出。数据集涵盖了三种常见的编辑任务,用于训练自监督的潜在扩散模型。通过该数据集,模型能够学习如何从示例对中推断出转换并将其应用于新的输入音频。该数据集的应用领域包括音频编辑和声音设计,旨在解决现有文本条件模型在处理复杂音频转换任务时的不足。
This dataset was developed by a research team at the University of California, Berkeley, specifically for audio texture manipulation tasks. It integrates speech data from LibriSpeech and VCTK, alongside environmental sound effects from BBC SFX. Each sample consists of an example input, an example output, a novel input audio clip, and a corresponding transformed output. The dataset includes three common editing tasks, designed for training self-supervised latent diffusion models. Using this dataset, models can learn to infer audio transformation rules from paired examples and apply these rules to new input audio clips. Its application domains cover audio editing and sound design, and it aims to address the shortcomings of current text-conditioned models when handling complex audio transformation tasks.




