ReasonAQA
收藏资源简介:
ReasonAQA数据集是由卡内基梅隆大学的研究团队创建的,旨在提升小规模音频语言模型在音频和文本上的推理能力。该数据集混合了现有数据集和合成数据,总共包含约56k个音频文件和1M个AQA实例,分为预训练、验证和测试三个部分。数据集来源于AudioCaps和Clotho,这两个数据集都包含了丰富的人类标注的音频描述。ReasonAQA的设计允许研究者在控制数据规模不变的情况下,研究模型设计、数据生成方法和预训练策略对推理性能的影响。
The ReasonAQA dataset was developed by a research team at Carnegie Mellon University, with the goal of enhancing the reasoning capabilities of small-scale audio language models across both audio and text modalities. This dataset combines existing datasets and synthetic data, comprising approximately 56,000 audio files and 1 million AQA instances, which are split into three subsets: pre-training, validation, and test. Sourced from AudioCaps and Clotho—two datasets containing extensive human-annotated audio descriptions—the ReasonAQA dataset enables researchers to investigate the effects of model design, data generation methods, and pre-training strategies on reasoning performance while maintaining a fixed data scale.

- 1Mellow: a small audio language model for reasoning卡内基梅隆大学 · 2025年



