遇见数据集

Early Screening of Prediabetes and Diabetes Using Multiclass Machine Learning Models

收藏
Zenodo2026-08-12 更新2026-08-13 收录
官方服务:

资源简介:

Diabetes is a serious global health issue, with the number of people living with diabetes continuing to rise year after year. Therefore, early identification of prediabetes and diabetes is essential to prevent more serious consequences. This research aims to develop a machine learning-based model to identify individuals with prediabetes and diabetes and to identify the factors that influence the model’s ability to identify diabetes cases. This research utilizes secondary data from the 2021 Behavioral Risk Factor Surveillance System (BRFSS), which contains 236,378 records. Data preprocessing includes handling data duplication, anomalies, and data balancing using the SMOTEENN resampling technique. The three machine learning algorithms used in this research are Random Forest, Support Vector Machine (SVM), and XGBoost, which were evaluated using accuracy, precision, recall, and F1 score metrics. The results show that XGBoost is the algorithm model with the best performance, achieving an F1 score of 0.96 for non-diabetes, 0.98 for prediabetes, and 0.96 for diabetes. A further result is the identification of ten key variables that influence the model’s decision-making in classifying non-diabetes, prediabetes, and diabetes, providing relevant information regarding patients’ characteristics and risk patterns across various health and behavioral aspects.

提供机构:
Zenodo
创建时间:
2026-08-12
二维码
社区交流群
二维码
科研交流群
商业服务