AudioMNIST
收藏资源简介:
AudioMNIST是由弗劳恩霍夫海因里希赫兹研究所人工智能部创建的一个公开音频数据集,包含30,000个英语口语数字的音频样本,总计约9.5小时的录音。该数据集用于语音数字和说话人性别分类任务,旨在为音频领域的模型架构和XAI算法提供基本的分类基准。数据集的创建过程包括使用RØDE NT-USB麦克风在安静的办公室环境中录制,保存为16位整数格式,并收集了包括年龄、性别、来源和口音在内的元信息。AudioMNIST的应用领域主要集中在自动语音识别和解释性人工智能的研究,旨在解决模型透明度和预测验证的问题。
AudioMNIST is a public audio dataset created by the Artificial Intelligence Division of the Fraunhofer Heinrich Hertz Institute. It contains 30,000 audio samples of spoken English digits, with a total recording duration of approximately 9.5 hours. This dataset is used for speech digit and speaker gender classification tasks, aiming to provide a basic classification benchmark for model architectures and XAI algorithms in the audio domain. The dataset's creation process involved recording in a quiet office environment using a RØDE NT-USB microphone, storing the audio in 16-bit integer format, and collecting metadata including age, gender, origin and accent. The main application fields of AudioMNIST focus on research in automatic speech recognition and explainable artificial intelligence, with the goal of addressing issues related to model transparency and prediction validation.

- 1AudioMNIST: Exploring Explainable Artificial Intelligence for Audio Analysis on a Simple Benchmark弗劳恩霍夫海因里希赫兹研究所人工智能部 · 2023年



