AIR-Bench
收藏资源简介:
AIR-Bench是首个针对大型音频-语言模型(LALMs)的生成评估基准,由浙江大学和阿里巴巴集团共同创建。该数据集包含约19000个单选题和2000个开放式问答数据,覆盖了人类语音、自然声音和音乐等多种音频类型。数据集通过创新的音频混合策略,如响度控制和时间错位,增强了音频的复杂性,更接近真实世界场景。AIR-Bench旨在全面评估LALMs在理解各种音频信号和遵循指令进行交互的能力,为未来研究提供方向和指导。
AIR-Bench is the first generative evaluation benchmark for Large Audio-Language Models (LALMs), co-developed by Zhejiang University and Alibaba Group. This dataset comprises approximately 19,000 multiple-choice questions and 2,000 open-ended question-answering samples, covering various audio types including human speech, natural sounds, and music. The dataset enhances audio complexity via innovative audio mixing strategies such as loudness control and temporal misalignment, making it more aligned with real-world scenarios. AIR-Bench aims to comprehensively evaluate the capabilities of LALMs in understanding diverse audio signals and following instructions for interaction, providing directions and guidance for future research.

- 1AIR-Bench: Automated Heterogeneous Information Retrieval Benchmark中国科学技术大学 · 2024年



