遇见数据集

adalat-ai/vividh-test-hindi

收藏
Hugging Face2026-05-20 更新2026-06-14 收录
官方服务:

资源简介:

Vividh-ASR基准测试(印地语测试集)是一个专门用于自动语音识别(ASR)模型评估的数据集,专注于印地语。它通过声学复杂性分层(分为Tier A到Tier D四个层级)来系统评估模型在真实世界音频中的表现,旨在揭示和解决模型在工作室录制语音上的偏见。数据集包含总计36.96小时、22,004个样本,来源包括FLEURS、IndicTTS、Kathbath、Common Voice、MUCS、Shrutilipi和IndicVoices等,覆盖了从清晰阅读语音到广播、自发语音和合成噪声等多种场景。该数据集仅用于测试,不用于训练,特别强调Tier D作为零样本压力测试,以评估模型的声学泛化能力。

The Vividh-ASR benchmark (Hindi test set) is a dataset dedicated to evaluating automatic speech recognition (ASR) models, with a primary focus on the Hindi language. It employs acoustic complexity stratification (four tiers from Tier A to Tier D) to systematically assess model performance on real-world audio, aiming to uncover and address biases in models toward studio-recorded speech. The dataset comprises a total of 36.96 hours of audio across 22,004 samples, sourced from resources including FLEURS, IndicTTS, Kathbath, Common Voice, MUCS, Shrutilipi, and IndicVoices, covering diverse scenarios ranging from clear read speech, broadcast content, spontaneous speech to audio with synthetic noise. This dataset is exclusively for testing purposes and not intended for model training, with Tier D specifically designated as a zero-shot stress test to evaluate the acoustic generalization capability of ASR models.

提供机构:
adalat-ai
二维码
社区交流群
二维码
科研交流群
商业服务