Arabic Parkinson's Disease Speech Dataset
收藏资源简介:
This dataset contains audio recordings collected to investigate acoustic biomarkers and machine learning models for detecting Parkinson’s Disease (PD) in native Arabic speakers. Designed to support voice pathology detection, acoustic feature extraction, and cross-linguistic validation studies, it serves as a benchmark resource for speech processing and clinical AI applications. Participant Breakdown: 40 native Arabic-speaking subjects, comprising 17 diagnosed Parkinson’s Disease patients and 23 age-matched Healthy Controls (HC). Speech Tasks: Includes sustained vowel phonations (/a/, /i/, /u/), reading of a standardized phonetically balanced Arabic text passage, and spontaneous speech recordings. Audio Format: MP3 format (.mp3, 44.1 kHz) to reduce file size and optimize download efficiency. Detailed Metadata & Methodology: Full clinical demographics, participant selection criteria, experimental setup, and acoustic feature extraction methodologies are detailed in the accompanying published paper. All recordings are fully anonymized to remove personal identifiable information (PII) and are distributed for open research under the Creative Commons Attribution 4.0 International (CC BY 4.0) license. Associated Journal Publication: Hassanat, A. B. A., et al. "Machine Learning-Based Detection of Parkinson's Disease From Arabic Speech: A Cross-Linguistic Validation Study." Journal of Central Nervous System Disease, vol. 18, 2026, pp. 1–12. DOI: 10.1177/11795735261448278



