KRAFTON/KMMAU
收藏资源简介:
KMMAU是一个韩语多模态音频理解基准,用于评估语音模型在多样化音频理解任务上的性能。该数据集涵盖9个子集,包括年龄估计、性别识别、说话者数量检测、事实提取、一般计数、职业识别、主题总结、词频计数和词序验证。基准构建自三个韩语语音数据集:Seoul Corpus、KMSAV和KSS,总共包含2,204个样本。每个样本包含音频文件、相关问题、参考答案以及元数据(如原始数据集标识和音频时间戳)。
KMMAU is a Korean multimodal audio understanding benchmark designed to evaluate the performance of speech models across a diverse range of audio understanding tasks. This dataset covers nine subsets, including age estimation, gender recognition, speaker count detection, fact extraction, general counting, occupation recognition, topic summarization, word frequency counting, and word order verification. The benchmark is built upon three Korean speech datasets: Seoul Corpus, KMSAV, and KSS, and contains a total of 2,204 samples. Each sample includes an audio file, relevant questions, reference answers, as well as metadata such as original dataset identifiers and audio timestamps.




