DeSTA-AQA5M
收藏资源简介:
DeSTA-AQA5M是一个大规模的音频指令调优数据集,包含了5百万个音频-提示-响应三元组。数据集由来自50个公开可用的音频数据集的7,000小时多样化的音频数据组成,涵盖了语音、环境声音和音乐等多种音频特征。该数据集通过自我生成的跨模态对齐策略构建,旨在解决大型音频语言模型(LALM)在任务无关的情况下实现鲁棒的听觉感知和指令遵循问题。数据集的创建过程利用了骨干语言模型生成自己的训练目标,从而在保留语言模型原有语言能力的同时,建立了有效的音频-文本对齐,实现了零样本泛化。
DeSTA-AQA5M is a large-scale audio instruction tuning dataset containing 5 million audio-prompt-response triplets. It consists of 7,000 hours of diverse audio data sourced from 50 publicly available audio datasets, covering various audio characteristics such as speech, environmental sounds, and music. Constructed via a self-generated cross-modal alignment strategy, this dataset aims to address the challenge of enabling large audio language models (LALMs) to achieve robust auditory perception and instruction following in a task-agnostic manner. Its creation process leverages a backbone language model to generate its own training objectives, thereby establishing effective audio-text alignment while preserving the original linguistic capabilities of the language model, and enabling zero-shot generalization.
DeSTA2.5-Audio数据集概述
基本信息
- 数据集名称:DeSTA2.5-Audio
- 托管地址:https://github.com/kehanlu/DeSTA2.5-Audio
数据描述
(注:根据提供的README内容,该数据集未包含具体描述信息)
使用说明
(注:根据提供的README内容,该数据集未包含使用说明)
其他信息
(注:根据提供的README内容,该数据集未提供其他相关信息)

- 1DeSTA2.5-Audio: Toward General-Purpose Large Audio Language Model with Self-Generated Cross-Modal Alignment清华大学 · 2025年



