Archit00/bbench-dep-air-bench
收藏资源简介:
AIR-Bench(音频指令基准)是首个专门设计用于评估大型音频-语言模型(LALMs)能力的基准测试。它旨在测试模型对多种音频信号(包括人类语音、自然声音和音乐)的理解能力,并进一步评估模型以文本形式与人类交互的能力。数据集包含两个维度:基础基准和聊天基准。基础基准由19个任务组成,包含约19,000个单项选择题;聊天基准则包含2,000个开放性问题-答案实例。该基准测试支持对LALMs在音频理解方面的全面评估,并提供了自动评分框架(使用GPT-4模型进行评分)。数据集适用于研究音频-语言模型性能比较和基准测试场景。
AIR-Bench (Audio InstRuction Benchmark) is the first benchmark designed to evaluate the ability of Large Audio-Language Models (LALMs) to understand various types of audio signals, including human speech, natural sounds, and music, and to interact with humans in textual format. The benchmark encompasses two dimensions: foundation benchmark and chat benchmark. The foundation benchmark consists of 19 tasks with approximately 19,000 single-choice questions, while the chat benchmark contains 2,000 instances of open-ended question-and-answer data. It provides an automated evaluation framework (using GPT-4 for scoring) for comprehensive assessment of LALMs in audio comprehension tasks, suitable for benchmarking and comparing model performances in audio-language domains.




