WordVoice-5A
收藏资源简介:
WordVoice-5A是由华南理工大学等机构联合构建的大规模双语细粒度标注数据集,旨在为基于大语言模型的文本到语音系统提供精确的词级声学控制基础。该数据集包含约4684小时的语音数据,涵盖中英文双语,总计约5226万个字符或词汇,并提供了时长、边界、能量、音高和语调五个维度的词级标注。数据创建过程采用严格的语音学指导流程,通过双重对齐模型提取时间戳,并结合响度优化和一致性检查,确保了标注的高精度与可靠性。该数据集主要应用于语音合成领域,特别支持有声书朗读和视频配音等需要高精度词级声学干预的场景,以解决现有TTS系统在细粒度控制方面的瓶颈问题。
WordVoice-5A is a large-scale bilingual fine-grained annotated dataset jointly constructed by South China University of Technology and other institutions. It aims to provide a precise word-level acoustic control foundation for text-to-speech (TTS) systems based on large language models (LLMs). This dataset contains approximately 4,684 hours of speech data covering both Chinese and English, with a total of about 52.26 million characters or words, and provides word-level annotations in five dimensions: duration, boundary, energy, pitch, and intonation. The construction of this dataset follows a strict phonetics-guided workflow, where timestamps are extracted via a dual alignment model, combined with loudness optimization and consistency checks to ensure high accuracy and reliability of the annotations. This dataset is mainly applied in the field of speech synthesis, and specifically supports scenarios requiring high-precision word-level acoustic intervention such as audiobook narration and video dubbing, aiming to address the bottleneck problem of existing TTS systems in fine-grained control.
数据集核心信息
- 数据集名称: WordVoice-5A
- 数据集规模: 4,700小时(4.7k-hour)双语数据
- 标注维度: 5维字级标注
- 时长 (Duration)
- 边界 (Boundary)
- 能量 (Energy)
- 音高 (Pitch)
- 音调 (Tone)
- 数据用途: 支持基于大语言模型(LLM)的高精度、多维度字级解耦控制语音合成(TTS),实现显式的字级韵律控制与零样本合成稳定性。
所属模型与框架
- 模型名称: WordVoice
- 核心机制:
- 边界符(Bound-Token)机制: 在LLM中实现自适应多任务韵律规划与灵活人工干预。
- 细粒度声学调制模块: 增强token到波形(token-to-waveform)阶段的控制能力。
- 合成模式:
- 自由模式(Free Mode): 模型自主进行韵律规划。
- 控制模式(Control Mode): 用户可显式操纵特定单词的五维声学属性。
- 1WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS华南理工大学; 虎牙公司; 同济大学; 香港理工大学; 佛山大学 · 2026年




