NeuroVoz: a Castillian Spanish corpus of parkinsonian speech
收藏资源简介:
The NeuroVoz dataset emerges as a pioneering resource in the field of computational linguistics and biomedical research, specifically designed to enhance the diagnosis and understanding of Parkinson's Disease (PD) through speech analysis. This dataset is distinguished as the first of its kind to be made publicly available in Castilian Spanish, addressing a critical gap in the availability of linguistic and dialectical diversity within PD research. Compiled from a cohort of 112 participants, including 54 individuals diagnosed with PD and 58 healthy controls, the NeuroVoz dataset offers a rich compilation of speech recordings. All PD participants were recorded under medication (ON state), ensuring consistency and reliability in the speech samples collected. The dataset is meticulously curated to include a variety of speech tasks—ranging from sustained vowel phonations and diadochokinetic (DDK) tests to 16 structured listen-and-repeat utterances and spontaneous monologues. The inclusion of both manually transcribed listen-and-repeat tasks and Whisper-automated transcriptions for monologues underscores our commitment to data accuracy and usability. Encompassing 2,977 audio files, the NeuroVoz dataset provides an extensive repository, averaging $26.88 \pm 3.35$ recordings per participant, making it an invaluable asset for researchers seeking to explore the nuances of PD-affected speech. The dataset's structure and composition facilitate a multifaceted analysis of speech impairments associated with PD, offering insights into phonatory, articulatory, and prosodic changes. In contributing to the body of knowledge with the NeuroVoz dataset, we invite the scientific community to engage with this dataset, explore the specific speech characteristics of PD in Castilian Spanish speakers, and advance the field of PD diagnosis through innovative speech analysis techniques. If you use this dataset, please cite both this Zenodo and the arXiv preprint: arXiv preprint: J. Mendes-Laureano, J. A. Gómez-García, A. Guerrero-López,E. Luque-Buzo, J. D. Arias-Londoño, F. J. Grandas-Pérez, and J. I. Godino-Llorente, “Neurovoz: a castillian spanish corpus of parkinsonian speech,” arXiv preprint arXiv:2403.02371 (2024). Link: https://arxiv.org/abs/2403.02371 Zenodo dataset: Mendes-Laureano, J., Gómez-García, J. A., Guerrero-López, A., Luque-Buzo, E., Arias-Londoño, J. D., Grandas-Pérez, F. J., & Godino Llorente, J. I. (2024). NeuroVoz: a Castillian Spanish corpus of parkinsonian speech (1.0.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.10777657
NeuroVoz数据集是计算语言学与生物医学研究领域的开创性资源,专为通过语音分析提升帕金森病(Parkinson's Disease, PD)的诊断与认知理解而设计。该数据集是首个以卡斯蒂利亚语(Castilian Spanish)公开发布的同类资源,填补了帕金森病研究中语言学与方言多样性数据匮乏的关键空白。 该数据集收纳了112名受试者的语音录音,其中54名经确诊为帕金森病患者,58名为健康对照者,涵盖了丰富的语音样本。所有帕金森病受试者均在药物开启状态(ON state)下进行录音,确保了所采集语音样本的一致性与可靠性。该数据集经过精心整理,涵盖了多种语音任务:从持续元音发音、轮替运动(diadochokinetic, DDK)测试,到16组结构化跟读语句与自发独白。数据集同时包含了跟读任务的人工转录文本与独白的Whisper自动转录文本,彰显了我们对数据准确性与易用性的承诺。 该数据集共包含2977个音频文件,每位受试者平均拥有26.88±3.35条录音,构建了一个规模庞大的语音库,为探索帕金森病患者语音特征差异的研究者提供了极为宝贵的资源。数据集的结构与组成支持对帕金森病相关语音障碍的多维度分析,可为发声、发音与韵律变化的研究提供见解。 我们发布NeuroVoz数据集以助力相关学术研究,诚挚邀请科学界使用该数据集,探索卡斯蒂利亚语使用者的帕金森病特异性语音特征,并通过创新的语音分析技术推动帕金森病诊断领域的发展。 若使用本数据集,请同时引用该Zenodo数据集与arXiv预印本: arXiv预印本:J. Mendes-Laureano、J. A. Gómez-García、A. Guerrero-López、E. Luque-Buzo、J. D. Arias-Londoño、F. J. Grandas-Pérez与J. I. Godino-Llorente,《NeuroVoz:帕金森病语音卡斯蒂利亚语语料库》,arXiv预印本arXiv:2403.02371(2024年)。 链接:https://arxiv.org/abs/2403.02371 Zenodo数据集:Mendes-Laureano, J., Gómez-García, J. A., Guerrero-López, A., Luque-Buzo, E., Arias-Londoño, J. D., Grandas-Pérez, F. J. & Godino Llorente, J. I. (2024). NeuroVoz:帕金森病语音卡斯蒂利亚语语料库(1.0.0)[数据集]. Zenodo. https://doi.org/10.5281/zenodo.10777657



