Urdu Conversational Speech Dataset
收藏资源简介:
Urdu Conversational Speech Dataset是由拉合尔管理科学大学创建的第一个用于评估乌尔都语自动语音识别(ASR)模型的会话语音数据集。该数据集包含471个音频记录,时长约1.3小时,涵盖了4名女性和6名男性的会话内容。数据集的创建过程包括三次转录以确保准确性,内容涉及巴基斯坦独立日、小组项目、斋月和开斋节等多样化主题。该数据集旨在解决乌尔都语ASR模型在会话环境中的性能评估问题,特别是在低资源语言处理中的应用。
The Urdu Conversational Speech Dataset, developed by the Lahore University of Management Sciences, is the first conversational speech dataset tailored for evaluating Urdu automatic speech recognition (ASR) models. It consists of 471 audio recordings with a total duration of approximately 1.3 hours, featuring conversational content from 4 female and 6 male speakers. To ensure annotation accuracy, the dataset underwent three rounds of transcription during its development, with the covered topics spanning diverse themes including Pakistan's Independence Day, group projects, Ramadan, and Eid al-Fitr. This dataset aims to fill the gap in performance evaluation of Urdu ASR models in conversational scenarios, particularly for applications in low-resource language processing.




