ArtificialAnalysis/big_bench_audio
收藏资源简介:
Artificial Analysis Big Bench Audio数据集是Big Bench Hard问题子集的音频版本,用于评估支持音频输入的模型的推理能力。数据集包含1000个音频记录,涵盖Big Bench Hard的四个类别:形式谬误、导航、对象计数和谎言网络。所有音频均为英语,使用23种不同的声音配置生成。数据集的结构包括类别、官方答案、文件名和ID四个字段。数据集的创建旨在为原生音频模型在推理任务上的基准测试提供支持,确保所选类别在音频设置中不会导致不公平的惩罚。音频生成过程使用了OpenAI、Microsoft Azure和Amazon的模型,并通过计算Levenshtein距离和人工审查来验证音频的准确性。数据集主要关注美国和英国口音,可能忽略了其他低资源语言和口音。
The Big Bench Audio dataset is an audio version of a subset of Big Bench Hard questions, designed to evaluate the reasoning capabilities of models that support audio input. The dataset includes 1000 audio recordings across four categories: Formal Fallacies, Navigate, Object Counting, and Web of Lies. The audio is synthetically generated in English using 23 voices from top providers. The dataset is structured with fields including category, official_answer, file_name, and id. The creation rationale focuses on benchmarking native audio models on reasoning tasks, avoiding categories that might unfairly penalize audio models. The source data is derived from Big Bench Hard, with modifications to ensure audio generation quality. The dataset also discusses potential biases, such as overfitting to English and US/UK accents.




