遇见数据集

adalat-ai/vividh-test-malayalam

收藏
Hugging Face2026-05-20 更新2026-06-14 收录
官方服务:

资源简介:

Vividh-ASR基准测试是一个针对马拉雅拉姆语的自动语音识别评估数据集,专门设计用于测试模型在真实世界音频中的性能。该数据集按声学复杂性分层,包含四个层级:Tier A(工作室朗读语音,来自FLEURS、IndicTTS等来源)、Tier B(广播快速语音,来自Shrutilipi)、Tier C(自发众包语音,来自IndicVoices和Common Voice)和Tier D(合成噪声语音,来自Kathbath Hard)。总时长为26.31小时,共15717个样本。该基准测试旨在暴露模型在干净、工作室录制语音上微调时产生的“工作室偏见”,帮助用户诊断模型在不同声学条件下的表现,特别是针对自发语音等真实部署场景。数据集仅包含测试分割,用于评估目的,不包含训练数据。

The Vividh-ASR Benchmark is an evaluation dataset for automatic speech recognition in Malayalam, designed to test model performance on real-world audio. It is complexity-stratified into four tiers: Tier A (studio, read speech from sources like FLEURS, IndicTTS), Tier B (broadcast, fast speech from Shrutilipi), Tier C (spontaneous, crowdsourced speech from IndicVoices and Common Voice), and Tier D (synthetic noise from Kathbath Hard). The total duration is 26.31 hours with 15,717 samples. This benchmark aims to expose studio-bias in models fine-tuned predominantly on clean, studio-recorded speech, helping diagnose model performance across different acoustic conditions, especially for real-world deployment scenarios like spontaneous speech. The dataset contains only the test split and is intended for evaluation purposes, with no training data included.

提供机构:
adalat-ai
二维码
社区交流群
二维码
科研交流群
商业服务