STARS
收藏资源简介:
STARS是一个用于歌唱转录、对齐和细化风格注释的统一框架,旨在解决歌唱语音合成(SVS)中对高质量标注数据集的需求。该框架提供多级标注,包括精确的音素-音频对齐、鲁棒的音符转录和时序定位、表现力的声乐技巧识别以及包括情感和节奏在内的全局风格特征。STARS通过分层声学特征处理,实现了帧、词、音素、音符和句子级别的多层次标注。该框架不仅克服了创建歌唱数据集的关键可扩展性挑战,还为可控的歌唱语音合成开辟了新的方法。
STARS is a unified framework for singing transcription, alignment, and refined style annotation, aiming to address the demand for high-quality annotated datasets in singing voice synthesis (SVS). This framework provides multi-level annotations, including accurate phoneme-audio alignment, robust note transcription and temporal localization, expressive vocal technique recognition, and global style features such as emotion and rhythm. STARS achieves multi-level annotations at frame, word, phoneme, note, and sentence levels through hierarchical acoustic feature processing. This framework not only overcomes the critical scalability challenges in constructing singing datasets, but also opens up new approaches for controllable singing voice synthesis.
STARS数据集概述
框架简介
- STARS是一个统一框架,同时解决歌唱转录、对齐和精细风格标注问题
- 提供多层次注释:
- 精确的音素-音频对齐
- 稳健的音符转录和时间定位
- 富有表现力的声乐技巧识别
- 全局风格特征(包括情感和节奏)
自动歌唱标注(ASA)示例
示例1
- 歌词:也 许 下 个 冬 天 <AP> 也 许 还 十 年
- 音素:ie x v x ia g e d ong t ian <AP> ie x v h ai sh i n ian
示例2
- 歌词:一 次 就 好 <AP> 我 带 你 去 看 天 荒 地 老
- 音素:i c i j iou h ao <AP> uo d ai n i q v k an t ian h uang d i l ao
示例3
- 歌词:my heads under water but <AP> im breathing fine <AP>
- 音素:MAY1 HH EH1 D Z AH1 N D ER0 W AA1 T ER0 B AH1 T <AP> AY1 M B R IY1 DH IH0 NG IH1 N F AY1 N <AP>
歌唱语音合成(SVS)应用
全局风格控制
- 音域:low, medium, high
- 节奏:slow, moderate, fast
- 情感:happy, sad
音素级技巧控制
- 可用技巧:mixed, falsetto, breathy, pharyngeal, vibrato, glissando, weak, strong, bubble
合成示例1
- 歌词:<SP> 不 再 看 天 上 太 阳 透 过 云 彩 的 光
- 全局风格:high, moderate, sad
- 音素技巧:详细标注每个音素的技巧编号(0-9)
合成示例2
- 歌词:在 阳 光 灿 烂 的 日 子 里 开 怀 大 笑
- 全局风格:medium, fast, happy
- 音素技巧:详细标注每个音素的技巧编号(0-9)
合成示例3
- 歌词:<SP> 远 处 蔚 蓝 天 空 下 涌 动 着 <AP> 金 色 的 麦 浪
- 全局风格:low, slow, happy
- 音素技巧:详细标注每个音素的技巧编号(0-9)




