b1n1yam/waxal-orm-tts-merged
收藏官方服务:
资源简介:
该数据集合并了来自google/WaxalNLP的人工标注奥罗莫语自动语音识别(ASR)分割和来自israel/waxal-autolabled的自动标注奥罗莫语分割。针对文本到语音(TTS)使用,自动标注转录中的前导[ORM]语言标签已被移除。数据行包含text和transcription字段,两者具有相同的清理后的值。数据集主要用于奥罗莫语的文本到语音和自动语音识别任务。
This dataset merges the human-labeled Oromo ASR split from google/WaxalNLP with the autolabeled Oromo split from israel/waxal-autolabled. For TTS use, the leading [ORM] language tag has been removed from autolabeled transcriptions. Rows include both text and transcription with the same cleaned value. The dataset is intended for Oromo text-to-speech and automatic speech recognition tasks.
提供机构:
b1n1yam


