ghananlpcommunity/twi-health-asr
收藏资源简介:
Twi健康语音数据集是一个领域特定的语音识别数据集,专注于加纳最广泛使用的语言之一——Twi,数据来源于公开的健康和保健主题视频内容。该数据集由Mich-Seth Owusu创建,并由GhanaNLP社区发布,旨在支持加纳语言的开源自动语音识别工具开发,特别涵盖健康和保健领域的词汇和话语。数据集包含75,441个音频片段,总时长约629小时,音频格式为WAV、16 kHz单声道,每个片段最长30秒。音频从公开的健康教育、传统医学、营养、健身和健康访谈节目等视频中提取,并分割成片段。转录使用Google Speech Recognition API(免费版)自动生成,但由于Twi是低资源语言且医学术语复杂,转录可能存在较高错误率,因此建议用户将其视为嘈杂的银标准标签,而非黄金标准真值。数据集结构包括音频列(采样率16000 Hz的音频)和转录列(Twi文本),仅提供训练分割。该数据集旨在填补Twi ASR数据的空白,作为社区迭代改进的基线资源,适用于预训练、弱监督和半监督学习等方法。
The Twi Healthcare Speech Dataset is a domain-specific automatic speech recognition (ASR) dataset focused on Twi, one of the most widely used languages in Ghana. The data is sourced from publicly available video content related to health and wellness topics. Created by Mich-Seth Owusu and released by the GhanaNLP community, this dataset aims to support the development of open-source automatic speech recognition tools for Ghanaian languages, with a specific focus on vocabulary and utterances in the healthcare and wellness domain. The dataset contains 75,441 audio clips with a total duration of approximately 629 hours. The audio is in WAV format, 16 kHz mono, with each clip lasting up to 30 seconds. The audio is extracted and segmented from publicly available videos covering health education, traditional medicine, nutrition, fitness, and health interview programs. Transcriptions were automatically generated using the Google Speech Recognition API (free tier). However, as Twi is a low-resource language and medical terminology is complex, the transcriptions may have a relatively high error rate. Users are therefore advised to treat these as noisy silver standard labels rather than gold standard ground truth. The dataset structure includes an audio column (audio sampled at 16000 Hz) and a transcription column (Twi text), with only the training split provided. This dataset aims to fill the gap in Twi ASR data, serving as a baseline resource for community-driven iterative improvements, and is applicable to methods such as pre-training, weakly-supervised learning, and semi-supervised learning.




