遇见数据集

anonyemnlpauthor18/KOpenAudioBench

收藏
Hugging Face2026-05-25 更新2026-05-31 收录
官方服务:

资源简介:

KOpenAudioBench是一个韩语口语问答基准测试数据集,源自OpenAudioBench。它是韩语语音基准测试套件的一部分,与KVoiceBench和KMMAU一起用于评估SpeechLMs。数据集包含2,835个韩语口语QA样本,分布在4个子集中:2,221个短答案问题和614个开放式提示。它强调短事实答案和开放式口语提示,涵盖历史/地理、娱乐/艺术、人文、实用知识、体育和自然科学等类别。数据集通过真实答案修正、超翻译、语音友好归一化和TTS合成构建,其中超翻译使用规则手册将源语言问题转换为韩语,并处理韩语特定重设计。由于许多项目需要短事实答案,答案别名处理是转换过程的一部分。

KOpenAudioBench is a Korean spoken question answering (SpokenQA) benchmark derived from OpenAudioBench. It is part of a Korean speech benchmark suite for evaluating SpeechLMs together with KVoiceBench and KMMAU. KOpenAudioBench contains 2,835 Korean spoken QA samples across 4 subsets: 2,221 short-answer questions and 614 open-ended prompts. It emphasizes short factual answers and open-ended spoken prompts covering categories such as history/geography, entertainment/arts, humanities, practical knowledge, sports, and natural science. KOpenAudioBench is constructed with the same human-agent SpokenQA benchmark transfer framework used for KVoiceBench, including ground-truth correction, hypertranslation, speech-friendly normalization, and TTS synthesis, where hypertranslation transfers source-language questions into Korean using a rulebook that handles Korean-specific redesign. Because many items require short factual answers, answer alias handling is part of the transfer process.

提供机构:
anonyemnlpauthor18
二维码
社区交流群
二维码
科研交流群
商业服务