遇见数据集

mrfakename/GLOBE_V2_Fixed

收藏
Hugging Face2025-12-07 更新2025-12-20 收录
官方服务:

资源简介:

--- license: cc0-1.0 --- # A version of the GLOBE dataset that works with `load_dataset` # Important notice Differences between V2 version and [the version described in paper](https://huggingface.co/datasets/MushanW/GLOBE): 1. The V2 version provide audio in 44.1kHz sample rate. (Supersampling) 2. The V2 versionn removed some samples (~5%) due to the volumn and text aligment issues. # Globe The full paper can be accessed here: [arXiv](https://arxiv.org/abs/2406.14875) An online demo can be accessed here: [Github](https://globecorpus.github.io/) ## Abstract This paper introduces GLOBE, a high-quality English corpus with worldwide accents, specifically designed to address the limitations of current zero-shot speaker adaptive Text-to-Speech (TTS) systems that exhibit poor generalizability in adapting to speakers with accents. Compared to commonly used English corpora, such as LibriTTS and VCTK, GLOBE is unique in its inclusion of utterances from 23,519 speakers and covers 164 accents worldwide, along with detailed metadata for these speakers. Compared to its original corpus, i.e., Common Voice, GLOBE significantly improves the quality of the speech data through rigorous filtering and enhancement processes, while also populating all missing speaker metadata. The final curated GLOBE corpus includes 535 hours of speech data at a 24 kHz sampling rate. Our benchmark results indicate that the speaker adaptive TTS model trained on the GLOBE corpus can synthesize speech with better speaker similarity and comparable naturalness than that trained on other popular corpora. We will release GLOBE publicly after acceptance. ## Citation ``` @misc{wang2024globe, title={GLOBE: A High-quality English Corpus with Global Accents for Zero-shot Speaker Adaptive Text-to-Speech}, author={Wenbin Wang and Yang Song and Sanjay Jha}, year={2024}, eprint={2406.14875}, archivePrefix={arXiv}, } ```

许可证:CC0 1.0 # 适配`load_dataset`的GLOBE数据集版本 ## 重要声明 V2版本与[论文中描述的版本](https://huggingface.co/datasets/MushanW/GLOBE)的差异如下: 1. V2版本提供44.1kHz采样率的音频(采用超采样处理) 2. V2版本移除了约5%因音量与文本对齐问题存在缺陷的样本 # GLOBE数据集 完整论文可通过以下链接获取:[arXiv](https://arxiv.org/abs/2406.14875) 在线演示可通过以下链接访问:[Github](https://globecorpus.github.io/) ## 摘要 本论文介绍了GLOBE——一款涵盖全球口音的高质量英语语料库,专为解决当前零样本(Zero-shot)说话人自适应文本转语音(Text-to-Speech, TTS)系统在适配带口音说话人时泛化能力不足的局限而设计。与LibriTTS、VCTK等主流英语语料库相比,GLOBE的独特优势在于其收录了来自23519位说话人的语音片段,覆盖全球164种口音,并为所有说话人配备了详细的元数据。相较于其原始语料库Common Voice,GLOBE通过严格的筛选与增强流程大幅提升了语音数据的质量,同时补全了所有缺失的说话人元数据。最终整理完成的GLOBE语料库包含535小时24kHz采样率的语音数据。我们的基准测试结果显示,基于GLOBE语料库训练的说话人自适应TTS模型,相较于基于其他主流语料库训练的同类模型,能够合成出与目标说话人特征匹配度更高、自然度相当的语音。论文被录用后,我们将公开发布GLOBE数据集。 ## 引用 @misc{wang2024globe, title={GLOBE: A High-quality English Corpus with Global Accents for Zero-shot Speaker Adaptive Text-to-Speech}, author={Wenbin Wang and Yang Song and Sanjay Jha}, year={2024}, eprint={2406.14875}, archivePrefix={arXiv}, }

提供机构:
mrfakename
二维码
社区交流群
二维码
科研交流群
商业服务