adalat-ai/vividh-test-malayalam
收藏资源简介:
Vividh-ASR基准测试是一个针对马拉雅拉姆语的自动语音识别评估数据集,专门设计用于测试模型在真实世界音频中的性能。该数据集按声学复杂性分层,包含四个层级:Tier A(工作室朗读语音,来自FLEURS、IndicTTS等来源)、Tier B(广播快速语音,来自Shrutilipi)、Tier C(自发众包语音,来自IndicVoices和Common Voice)和Tier D(合成噪声语音,来自Kathbath Hard)。总时长为26.31小时,共15717个样本。该基准测试旨在暴露模型在干净、工作室录制语音上微调时产生的“工作室偏见”,帮助用户诊断模型在不同声学条件下的表现,特别是针对自发语音等真实部署场景。数据集仅包含测试分割,用于评估目的,不包含训练数据。
The Vividh-ASR Benchmark is an evaluation dataset for automatic speech recognition in Malayalam, designed to test model performance on real-world audio. It is complexity-stratified into four tiers: Tier A (studio, read speech from sources like FLEURS, IndicTTS), Tier B (broadcast, fast speech from Shrutilipi), Tier C (spontaneous, crowdsourced speech from IndicVoices and Common Voice), and Tier D (synthetic noise from Kathbath Hard). The total duration is 26.31 hours with 15,717 samples. This benchmark aims to expose studio-bias in models fine-tuned predominantly on clean, studio-recorded speech, helping diagnose model performance across different acoustic conditions, especially for real-world deployment scenarios like spontaneous speech. The dataset contains only the test split and is intended for evaluation purposes, with no training data included.



