遇见数据集

zuhri025/ultra-unified-easy

收藏
Hugging Face2026-05-07 更新2026-05-31 收录
官方服务:

资源简介:

Ultra Unified Easy是一个专为稳定文本转语音(TTS)训练阶段设计的预处理数据集,包含英语和乌尔都语的双语短序列数据。数据集从原始数据集humair025/ultra-unified中过滤生成,适用于课程学习(curriculum learning)方法。数据规则限制每个序列的token长度小于150,确保训练稳定性。数据集包含919,314条训练数据,字段包括文本内容、音素标注、内容token索引、全局嵌入向量、token长度及数据来源。该数据集通过结构化字段支持TTS模型的高效训练。

Ultra Unified Easy is a preprocessed dataset designed for stable text-to-speech (TTS) training phases, containing bilingual short sequences in English and Urdu. It is filtered from the original dataset humair025/ultra-unified for curriculum learning purposes. The dataset enforces a rule where token length is less than 150 to ensure training stability. It comprises 919,314 training rows with fields including text, phonemes, content token indices, global embedding, token length, and source. The structured fields facilitate efficient TTS model training.

提供机构:
zuhri025
二维码
社区交流群
二维码
科研交流群
商业服务