MMAR
收藏资源简介:
MMAR是一个新的基准,旨在评估音频语言模型(ALM)在大量多学科任务中的深度推理能力。MMAR由1000个精心策划的音频-问答三元组组成,从现实世界的互联网视频中收集,并通过迭代错误校正和质量检查进行精炼,以确保高质量。与现有仅限于声音、音乐或特定领域语音的基准不同,MMAR将它们扩展到广泛的真实世界音频场景,包括声音、音乐和语音的混合模式组合。MMAR中的每个问题都按四个推理层进行分层分类:信号、感知、语义和文化,每个层中还有额外的子类别,以反映任务的多样性和复杂性。为了进一步促进该领域的研究,我们用思维链(CoT)推理为每个问题进行注释,以促进未来在音频推理方面的进步。基准中的每个项目都要求进行多步深度推理,超越表面理解。此外,部分问题需要研究生水平的感知和特定领域的知识,从而提高了基准的难度和深度。我们使用广泛的模型评估了MMAR,包括大型音频语言模型(LALM)、大型音频推理模型(LARM)、全能语言模型(OLM)、大型语言模型(LLM)和大型推理模型(LRM),并使用音频标题输入。这些模型在MMAR上的性能突显了基准的挑战性,我们的分析进一步揭示了当前模型在理解和推理能力方面的关键局限性。这些发现强调了在音频语言推理方面进行更多研究的紧迫性,包括数据和算法创新。我们希望MMAR将成为未来在这个重要但探索较少的领域取得进展的催化剂。
MMAR is a novel benchmark designed to evaluate the deep reasoning capabilities of Audio Language Models (ALMs) across a wide range of multidisciplinary tasks. MMAR consists of 1,000 carefully curated audio-QA triples collected from real-world internet videos, and refined through iterative error correction and quality assurance to ensure high quality. In contrast to existing benchmarks limited to sound, music, or domain-specific speech, MMAR expands their scopes to a broad array of real-world audio scenarios, including mixed combinations of sound, music, and speech modalities. Each question in MMAR is hierarchically categorized into four reasoning layers: signal, perceptual, semantic, and cultural, with additional subcategories under each layer to reflect the diversity and complexity of the tasks. To further promote research in this field, we annotated every question with Chain-of-Thought (CoT) reasoning to facilitate future advancements in audio-based reasoning. Each item in the benchmark requires multi-step deep reasoning that surpasses surface-level comprehension. Additionally, some questions demand graduate-level perceptual and domain-specific knowledge, thereby elevating the benchmark's difficulty and depth. We evaluated MMAR using a diverse set of models, including Large Audio Language Models (LALMs), Large Audio Reasoning Models (LARMs), Omnipotent Language Models (OLMs), Large Language Models (LLMs), and Large Reasoning Models (LRMs), with audio caption inputs. The performance of these models on MMAR highlights the benchmark's challenging nature, and our analysis further uncovers critical limitations in current models' comprehension and reasoning capabilities. These findings underscore the urgency of further research in audio-language reasoning, including data and algorithmic innovations. We hope that MMAR will serve as a catalyst for future progress in this important yet under-explored field.

- 1MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix上海交通大学、南洋理工大学、伦敦玛丽女王大学、字节跳动、上海创新研究院、清华大学、中国科学院大学、2023AI、香港科技大学、德克萨斯大学奥斯汀分校 · 2025年



