遇见数据集

zuhri025/ultra-unified-hard

收藏
Hugging Face2026-05-07 更新2026-05-31 收录
官方服务:

资源简介:

Ultra Unified Hard 是一个专为文本到语音(TTS)训练设计的数据集,包含长且复杂的序列,旨在提高TTS模型的鲁棒性。该数据集从humair025/ultra-unified中过滤而来,遵循token长度大于等于150的规则,确保数据具有挑战性。它支持英语和乌尔都语,并包含文本、音素、内容token索引、全局嵌入、token长度和来源等字段,适用于HARD课程阶段的高级训练。

Ultra Unified Hard is a dataset specifically designed for text-to-speech (TTS) training, featuring long and complex sequences to enhance the robustness of TTS models. It is filtered from humair025/ultra-unified with a rule that token length must be at least 150, ensuring challenging data. The dataset supports English and Urdu, and includes fields such as text, phonemes, content token indices, global embedding, token length, and source, making it suitable for advanced training in the HARD curriculum stage.

提供机构:
zuhri025
二维码
社区交流群
二维码
科研交流群
商业服务