Parkinsons Disease Speech Dataset
收藏资源简介:
该数据集来自牛津大学,包含195个实例,其中147个为帕金森病患者,48个为非患者。数据集包含22个特征,如频率、音高、声波的振幅/周期等,以及一个标签,1代表帕金森病,0代表非帕金森病。
This dataset originates from the University of Oxford and comprises 195 instances, including 147 patients with Parkinson's disease and 48 non-patients. The dataset encompasses 22 features, such as frequency, pitch, amplitude/period of sound waves, among others, along with a label where 1 denotes Parkinson's disease and 0 indicates non-Parkinson's disease.
数据集概述
数据集名称
A Machine Learning Approach for the Diagnosis of Parkinsons Disease via Speech Analysis
研究时间
March 2022
数据集来源
University of Oxford
数据集组成
- 实例数量: 195
- 147 Parkinsons subjects
- 48 without Parkinsons
- 特征数量: 22
- 包括频率、音高、声波的振幅/周期等特征
- 标签: 1代表Parkinson’s, 0代表无Parkinson’s
使用算法
- Logistic Regression (LR)
- Linear Discriminant Analysis (LDA)
- k Nearest Neighbors (KNN)
- Decision Tree (DT)
- Neural Network (NN)
- Naive Bayes (NB)
- Gradient Boost (GB)
工程目标
开发一个机器学习模型,用于Parkinson’s的诊断,至少达到90%的准确率和/或Matthews Correlation Coefficient至少为0.9。
数据分析结果
模型在数据集重新调整后,使用75-25的训练-测试分割表现最佳。K Nearest Neighbors和Neural Network达到了98%的最高准确率。
结论
该项目证明了机器学习在Parkinson’s诊断中相较于当前方法有显著改进,模型达到了98%的准确率,对于有效治疗至关重要。




