VCTK-Token
收藏资源简介:
VCTK-Token数据集由加州大学伯克利分校和加州大学旧金山分校的研究团队创建,旨在用于语音不流畅性检测。该数据集包含58405条语音和对应的标注文本,涵盖了多种不流畅性类型,如重复、插入、删除等。数据集的创建过程包括文本模拟器和语音模拟器的使用,生成带有不流畅标记的语音样本。该数据集主要应用于对话系统中的意图理解、语音治疗和语音障碍筛查等领域,旨在提高对语音不流畅性的检测和分类能力。
The VCTK-Token dataset was developed by a research team from the University of California, Berkeley, and the University of California, San Francisco, and is intended for speech disfluency detection. This dataset contains 58,405 speech samples and their corresponding annotated texts, covering various types of speech disfluencies such as repetitions, insertions, deletions, and so on. The development process of the dataset involves the use of text simulators and speech simulators to generate speech samples with disfluency markers. It is mainly applied in fields such as intent understanding in dialogue systems, speech therapy, and speech disorder screening, aiming to improve the detection and classification capabilities of speech disfluencies.

- 1Time and Tokens: Benchmarking End-to-End Speech Dysfluency Detection浙江大学, 加州大学伯克利分校, 加州大学旧金山分校 · 2024年



