遇见数据集

adalat-ai/fleurs-ro

收藏
Hugging Face2026-05-20 更新2026-06-21 收录
官方服务:

资源简介:

FLEURS-RO是一个测试专用的印度语言丰富转录基准数据集,源自google/fleurs数据集,用于自动语音识别任务。它支持三种印度语言:马来语(ml)、印地语(hi)和卡纳达语(kn),每种语言仅提供测试分割,分别包含957、418和838个测试样本。数据集特征包括音频(16 kHz单声道)、原始转录(逐字转录)、丰富转录(带语法标点、标准化数字格式和印度文字正字法修正)、持续时间(秒)和语言标签。该数据集通过LLM(Gemini 2.5 Pro)流程生成丰富转录目标,经过母语语言学家迭代优化提示,以改进标点、数字格式和正字法。它适用于评估语音识别模型,特别是与SCRIBE评估框架结合使用,进行诊断性分类评估(如WER、NER、PER、TER)。数据集遵循CC-BY-4.0许可证,基于上游FLEURS数据集。

FLEURS-RO is a test-only benchmark dataset for rich transcription of Indian languages, derived from the google/fleurs dataset and tailored for automatic speech recognition (ASR) tasks. It supports three Indian languages: Malayalam (ml), Hindi (hi), and Kannada (kn). Only the test split is provided for each language, containing 957, 418, and 838 test samples respectively. Dataset features include audio recordings (16 kHz, mono), raw verbatim transcriptions, rich transcriptions (with grammatical punctuation, standardized numeric formatting, and orthography corrections for Indian languages), duration (in seconds), and language labels. The rich transcription targets are generated using an LLM (Gemini 2.5 Pro) pipeline, with prompts iteratively refined by native-language linguists to enhance punctuation, numeric formatting, and orthographic accuracy. This dataset is suitable for evaluating speech recognition models, particularly for diagnostic classification assessments (such as WER, NER, PER, TER) when used in conjunction with the SCRIBE evaluation framework. It is released under the CC-BY-4.0 license and built upon the upstream FLEURS dataset.

提供机构:
adalat-ai
二维码
社区交流群
二维码
科研交流群
商业服务