AHELM
收藏资源简介:
AHELM是一个全面的音频-语言模型评估基准,旨在从技术和社会角度全面评估音频-语言模型的能力。它包括10个关键方面:音频感知、知识、推理、情感检测、偏见、公平性、多语言性、鲁棒性、毒性和安全性。AHELM汇集了各种数据集,包括2个新的合成音频-文本数据集PARADE和CoRe-Bench,以全面衡量音频-语言模型在10个重要方面的性能。此外,AHELM还标准化了提示、推理参数和评估指标,以确保模型之间的公平比较。
AHELM is a comprehensive audio-language model evaluation benchmark designed to comprehensively assess the capabilities of audio-language models from both technical and societal perspectives. It encompasses 10 key dimensions: audio perception, knowledge, reasoning, emotion detection, bias, fairness, multilingualism, robustness, toxicity, and safety. AHELM aggregates a diverse range of datasets, including two newly synthesized audio-text datasets PARADE and CoRe-Bench, to thoroughly evaluate the performance of audio-language models across these 10 critical dimensions. Furthermore, AHELM standardizes prompts, inference parameters, and evaluation metrics to enable fair and consistent comparisons between different models.
Holistic Evaluation of Audio-Language Models (AHELM) 数据集概述
数据集简介
AHELM 是一个用于全面评估音频-语言模型(ALMs)性能的基准测试。ALMs 是多模态模型,能够接收交错的音频和文本作为输入,并输出文本。
核心目标
解决现有评估中缺乏标准化基准的问题,通过聚合多个数据集,全面衡量 ALMs 在 10 个关键方面的表现。
评估维度
- 音频感知
- 知识
- 推理
- 情感检测
- 偏见
- 公平性
- 多语言性
- 鲁棒性
- 毒性
- 安全性
包含数据集
- PARADE:新的合成音频-文本数据集,评估 ALMs 避免刻板印象的能力。
- CoRe-Bench:新的合成音频-文本数据集,通过多轮问答对话音频测量推理能力。
标准化措施
- 统一提示词
- 统一推理参数
- 统一评估指标
官方资源
- 论文:https://crfm.stanford.edu/helm/audio/v1.0.0/
- GitHub:https://crfm.stanford.edu/helm/audio/v1.0.0/
- Leaderboard:https://crfm.stanford.edu/helm/audio/v1.0.0/




