VietAudio-team/Vietnamese-asr-leaderboard
收藏资源简介:
越南语开放ASR评估数据集是一个用于评估自动语音识别(ASR)模型性能的标准化基准数据集。它整合了9个最大的公开越南语语音数据集,包括vivos、common_voice、vivoice、vimd、lsvsc、vlsp、fosd、gigaspeech2和bud500,总计157,755个音频样本,总时长约212小时。数据集覆盖合成和真实环境录音,采样率多样(16kHz至48kHz)。所有文本转录均经过严格的标准化处理,包括字符和标点规范化(仅保留逗号、句号、问号和感叹号)以及数字到文本的转换(如将“20”转换为“hai mươi”)。数据通过合并现有测试/验证集或随机洗牌(种子42)策略构建,确保评估的客观性和可重复性。该数据集专为越南语ASR排行榜设计,提供统一的ground truth用于词错误率(WER)计算。
The Vietnamese Open ASR Evaluation Dataset is a standardized benchmark dataset for evaluating the performance of automatic speech recognition (ASR) models. It integrates 9 of the largest publicly available Vietnamese speech datasets, including vivos, common_voice, vivoice, vimd, lsvsc, vlsp, fosd, gigaspeech2, and bud500, totaling 157,755 audio samples with a duration of approximately 212 hours. The dataset covers both synthetic and real-world recordings with diverse sampling rates (16kHz to 48kHz). All text transcriptions undergo rigorous normalization, including character and punctuation standardization (retaining only commas, periods, question marks, and exclamation marks) and number-to-text conversion (e.g., converting 20 to hai mươi). Data is constructed by merging existing test/validation splits or through random shuffling (seed 42) to ensure objectivity and reproducibility in evaluation. This dataset is designed for the Vietnamese ASR Leaderboard, providing unified ground truth for word error rate (WER) calculation.



