遇见数据集

b1n1yam/waxal-orm-tts-merged

收藏
Hugging Face2026-05-15 更新2026-05-31 收录
官方服务:

资源简介:

该数据集合并了来自google/WaxalNLP的人工标注奥罗莫语自动语音识别(ASR)分割和来自israel/waxal-autolabled的自动标注奥罗莫语分割。针对文本到语音(TTS)使用,自动标注转录中的前导[ORM]语言标签已被移除。数据行包含text和transcription字段,两者具有相同的清理后的值。数据集主要用于奥罗莫语的文本到语音和自动语音识别任务。

This dataset merges the human-labeled Oromo ASR split from google/WaxalNLP with the autolabeled Oromo split from israel/waxal-autolabled. For TTS use, the leading [ORM] language tag has been removed from autolabeled transcriptions. Rows include both text and transcription with the same cleaned value. The dataset is intended for Oromo text-to-speech and automatic speech recognition tasks.

提供机构:
b1n1yam
二维码
社区交流群
二维码
科研交流群
商业服务