large-scale in-the-wild Japanese laughter corpus
收藏资源简介:
我们提供了一个大规模的野外日语笑声语料库和一种笑声合成方法。此前的笑声合成工作不仅缺乏数据,也缺乏适当的笑声表示方法。为了解决这些问题,我们首先提出了一个包含3.5小时笑声的野外语料库,据我们所知,这是为笑声合成设计的最庞大的笑声语料库。然后,我们提出了伪音素令牌(PPTs)来通过一系列离散令牌表示笑声,这些令牌是通过在从笑声中提取的特征上训练一个预训练的自监督模型上的聚类模型获得的。笑声可以通过将PPTs输入到文本到语音系统中来合成。我们还展示了PPTs可以用于训练一个语言模型,用于无条件笑声生成。综合主观和客观评估的结果表明,所提出的方法显著优于基线方法,并且可以无条件生成自然的笑声。
We present a large-scale in-the-wild Japanese laughter corpus and a method for laughter synthesis. Previous efforts in laughter synthesis have been hindered not only by a lack of data but also by the absence of appropriate methods for representing laughter. To address these issues, we first introduce a corpus containing 3.5 hours of in-the-wild laughter, which, to our knowledge, is the most extensive laughter corpus designed for laughter synthesis. Subsequently, we propose pseudo-phoneme tokens (PPTs) to represent laughter through a series of discrete tokens, obtained by training a clustering model on features extracted from laughter using a pre-trained self-supervised model. Laughter can be synthesized by inputting PPTs into a text-to-speech system. We also demonstrate that PPTs can be utilized to train a language model for unconditional laughter generation. Comprehensive subjective and objective evaluations indicate that the proposed method significantly outperforms baseline approaches and is capable of generating natural laughter unconditionally.
数据集概述
数据集名称
- Laughter Corpus
数据集描述
- 该数据集是一个大规模的日语笑声语料库,包含约3.5小时的笑声数据,是目前为止用于笑声合成的最大笑声语料库。
数据集用途
- 用于笑声合成研究,特别是结合伪音素标记(PPTs)进行笑声的合成和无条件生成。
数据集特点
- 通过预训练的自监督模型提取特征,并使用聚类模型生成伪音素标记(PPTs)来表示笑声。
- 支持通过文本到语音系统进行笑声合成,并可用于训练语言模型进行无条件笑声生成。
数据集评估
- 通过综合的主观和客观评估,证明该方法在笑声合成上显著优于基线方法,能够生成自然的无条件笑声。
数据集获取
- 数据集可通过以下链接下载:Laughter Corpus
数据集使用方法
- 数据集需下载并放置在指定目录下,通过预处理脚本进行数据准备,然后用于训练TTS模型和语言模型。
数据集引用信息
-
若使用此数据集,请引用以下论文:
@inproceedings{xin2023laughter title={Laughter Synthesis using Pseudo Phonetic Tokens with a Large-scale In-the-wild Laughter Corpus}, author={Xin, Detai and Takamichi, Shinnosuke and Morimatsu, Ai and Saruwatari, Hiroshi}, booktitle={Proc. Interspeech}, year={2023} }




