遇见数据集

AISHELL-Stammertalk 中文口吃数据库 A Mandarin stuttered speech dataset

收藏
Zenodo2024-07-13 更新2026-05-26 收录
官方服务:

资源简介:

Dataset official website: https://aishelltech.com/aishell_6AThis Zenodo page contains dataset samples. To access and download the full dataset, please send an application here https://opendata.aishelltech.com/stammertalk The AISHELL-Stammertalk datasets consists of recordings from 70 native mardarin AWS (Adults who stutter), including 46 males and 24 females. The total duration is 48.8 hours. Each participant engaged in a recording session lasting up to one hour, comprising two parts: conversation and voice command reading. Conversations were conducted through online interviews using platforms like Zoom or Tencent Meet, aiming to capture spontaneous speech on diverse topics. The interviewer, one of the two authors, posed questions based on a prepared list, with the flexibility to introduce impromptu questions as needed. In the voice command reading part, participants were tasked with reading a set of 200 commands, categorized into car navigation and smart home device interaction. To ensure variety, a new set of 200 commands was introduced for every 25 participants, resulting in a dataset featuring a total of 600 unique commands. Participants were encouraged to employ the Voluntary Stuttering technique, deliberately introducing stuttering. Five types of stuttering were specified by the annotation guidelines, including:[]: Word/phrase repetition. Designated for marking entire repeated character or phrase./b: block. Gasps for air or stuttered pauses./p: prolongation. Elongated phoneme./r: sound repetition. Repeated phoneme that do not constitute an entire character./i: interjections. Filler characters due to stuttering e.g., ‘嗯’, ‘啊’, or ‘呃’. Notably, naturally occurring interjections that don't disrupt the speech flow are excluded.

数据集官方网站:https://aishelltech.com/aishell_6A。本Zenodo页面仅包含数据集样本。如需访问并下载完整数据集,请通过以下链接提交申请:https://opendata.aishelltech.com/stammertalk。 AISHELL-Stammertalk数据集收录了70名母语为普通话的成人口吃者(Adults who Stutter, AWS)的语音录制数据,其中男性46名,女性24名,总录制时长为48.8小时。每位参与者的录制会话最长可达1小时,包含两个环节:对话访谈与语音指令朗读。对话环节通过Zoom、腾讯会议等在线访谈平台开展,旨在采集多主题的自发口语内容。访谈者为两位作者之一,将基于预设清单提出问题,并可根据实际需求灵活引入即兴问题。 在语音指令朗读环节,参与者需朗读200条语音指令,指令分为车载导航与智能家居设备交互两大类。为保证指令内容的多样性,每25名参与者将更换一组全新的200条指令,最终数据集共包含600条独特语音指令。参与者被鼓励使用自主口吃(Voluntary Stuttering)技术,刻意制造口吃现象。 标注指南明确规定了五类口吃类型:[标记]:词语/短语重复,用于标记完整重复的字符或短语;/b:阻塞,指换气困难或口吃性停顿;/p:拖音,即拉长的音素;/r:音素重复,指非完整字符的重复音素;/i:插入语,指因口吃产生的填充词,例如“嗯”“啊”或“呃”。需注意,未破坏言语流畅性的自然插入语将被排除。

提供机构:
Zenodo
创建时间:
2024-07-13
二维码
社区交流群
二维码
科研交流群
商业服务