AUDIOMARATHON
收藏资源简介:
AUDIOMARATHON是一个全面的声音理解基准,专为评估长上下文音频理解和推理效率而设计。该数据集由来自三个领域的多样化任务组成:语音、声音和音乐,以及覆盖十个代表性子任务的全面任务覆盖,包括语音识别、语音内容推理、语音实体识别、音乐分类、音频场景分类、声音事件检测、情绪识别、语音检测、说话人年龄识别和说话人性别识别。AUDIOMARATHON通过一个严格的多阶段框架构建,确保了多样性、难度和高标注质量,旨在推动音频和多模态研究社区开发更先进的音频理解模型,能够解决复杂的音频问题。
AUDIOMARATHON is a comprehensive audio understanding benchmark specifically designed to evaluate long-context audio understanding and reasoning efficiency. This dataset comprises diverse tasks across three domains: speech, sound, and music, and features comprehensive task coverage spanning ten representative subtasks including speech recognition, speech content reasoning, speech entity recognition, music classification, audio scene classification, sound event detection, emotion recognition, speech detection, speaker age recognition, and speaker gender recognition. AUDIOMARATHON is constructed via a rigorous multi-stage framework to ensure diversity, difficulty, and high annotation quality, aiming to promote the audio and multimodal research communities to develop more advanced audio understanding models capable of solving complex audio problems.




