SaSLaW
收藏资源简介:
SaSLaW是由东京大学等机构创建的自发性对话语音语料库,包含同步记录的说话者、听者和观看者的音频-视觉信息。该数据集旨在模拟真实环境中的语音通信,通过记录两位参与者在模拟嘈杂环境中的对话来收集数据。数据集的创建过程包括使用高采样率的麦克风和头戴式摄像机进行同步记录,以及对环境噪音的模拟。SaSLaW主要用于开发和评估环境适应性文本到语音合成模型,以实现更自然和无缝的对话通信。
SaSLaW is a spontaneous conversational speech corpus developed by The University of Tokyo and other institutions, containing synchronously recorded audio-visual information of speakers, listeners and observers. This corpus aims to simulate real-world voice communication, with data collected by recording dialogues between two participants in a simulated noisy environment. The construction of SaSLaW involves synchronous recording using high-sampling-rate microphones and head-mounted cameras, as well as simulation of environmental background noise. SaSLaW is primarily used to develop and evaluate environment-adaptive text-to-speech synthesis models, aiming to enable more natural and seamless conversational communication.




