遇见数据集

NeuroVoz: a Castillian Spanish corpus of parkinsonian speech

收藏
Zenodo2025-12-04 更新2026-05-26 收录
官方服务:

资源简介:

The NeuroVoz dataset emerges as a pioneering resource in the field of computational linguistics and biomedical research, specifically designed to enhance the diagnosis and understanding of Parkinson's Disease (PD) through speech analysis. This dataset is distinguished as the first of its kind to be made publicly available in Castilian Spanish, addressing a critical gap in the availability of linguistic and dialectical diversity within PD research. Compiled from a cohort of 112 participants, including 54 individuals diagnosed with PD and 58 healthy controls, the NeuroVoz dataset offers a rich compilation of speech recordings. All PD participants were recorded under medication (ON state), ensuring consistency and reliability in the speech samples collected. The dataset is meticulously curated to include a variety of speech tasks—ranging from sustained vowel phonations and diadochokinetic (DDK) tests to 16 structured listen-and-repeat utterances and spontaneous monologues. The inclusion of both manually transcribed listen-and-repeat tasks and Whisper-automated transcriptions for monologues underscores our commitment to data accuracy and usability. Encompassing 2,977 audio files, the NeuroVoz dataset provides an extensive repository, averaging 26.88 +- 3.35 recordings per participant, making it an invaluable asset for researchers seeking to explore the nuances of PD-affected speech. The dataset's structure and composition facilitate a multifaceted analysis of speech impairments associated with PD, offering insights into phonatory, articulatory, and prosodic changes. In contributing to the body of knowledge with the NeuroVoz dataset, we invite the scientific community to engage with this dataset, explore the specific speech characteristics of PD in Castilian Spanish speakers, and advance the field of PD diagnosis through innovative speech analysis techniques. If you use this dataset, please cite both this Zenodo and the article describing the corpus: Mendes-Laureano, J., Gómez-García, J.A., Guerrero-López, A. et al. NeuroVoz: a Castillian Spanish corpus of parkinsonian speech. Sci Data 11, 1367 (2024). https://doi.org/10.1038/s41597-024-04186-z Zenodo dataset: Mendes-Laureano, J., Gómez-García, J. A., Guerrero-López, A., Luque-Buzo, E., Arias-Londoño, J. D., Grandas-Pérez, F. J., & Godino Llorente, J. I. (2024). NeuroVoz: a Castillian Spanish corpus of parkinsonian speech (1.0.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.10777657

NeuroVoz数据集是计算语言学与生物医学研究领域的开创性资源,专为通过语音分析提升帕金森病(Parkinson's Disease, PD)的诊断与认知而设计。本数据集作为首个公开可用的卡斯蒂利亚西班牙语语料库,填补了帕金森病研究中语言学与方言多样性数据匮乏的关键空白。 该数据集源自112名受试者队列,其中54名经确诊为帕金森病患者,58名为健康对照者,包含丰富的语音录音样本。所有帕金森病受试者均在服药状态(ON期)下进行录音,确保所采集语音样本的一致性与可靠性。数据集经精心编纂,涵盖多种语音任务:包括持续元音发声、构音交替运动(diadochokinetic, DDK)测试、16组结构化跟读语句,以及自发性独白。数据集同时包含手动转录的跟读语句文本与基于Whisper自动生成的独白转录文本,彰显了我们对数据准确性与易用性的承诺。 NeuroVoz数据集总计包含2977个音频文件,平均每名受试者拥有26.88±3.35条录音,是探索帕金森病患者语音特征的宝贵研究资产。该数据集的结构与组成支持对帕金森病相关语音损伤的多维度分析,可用于揭示发声、发音与韵律层面的变化规律。 本数据集旨在为相关研究领域贡献知识储备,我们诚挚邀请科学界使用该数据集,探究卡斯蒂利亚西班牙语使用者的帕金森病特异性语音特征,并通过创新性语音分析技术推动帕金森病诊断领域的发展。 若使用本数据集,请同时引用本Zenodo存档与介绍该语料库的学术论文: Mendes-Laureano, J., Gómez-García, J.A., Guerrero-López, A. 等. NeuroVoz: 帕金森病语音卡斯蒂利亚西班牙语语料库. Sci Data 11, 1367 (2024). https://doi.org/10.1038/s41597-024-04186-z Zenodo数据集存档:Mendes-Laureano, J., Gómez-García, J. A., Guerrero-López, A., Luque-Buzo, E., Arias-Londoño, J. D., Grandas-Pérez, F. J., & Godino Llorente, J. I. (2024). NeuroVoz: a Castillian Spanish corpus of parkinsonian speech (1.0.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.10777657

提供机构:
Zenodo
创建时间:
2024-09-03
二维码
社区交流群
二维码
科研交流群
商业服务