Summary of data preprocessing steps.

Figshare2026-02-09 更新2026-04-28 收录

下载链接：

https://figshare.com/articles/dataset/_p_Summary_of_data_preprocessing_steps_p_/31300586

下载链接

链接失效反馈

官方服务：

资源简介：

Cardiovascular diseases (CVDs) are leading causes of morbidity and mortality globally, with a growing burden in low- and middle-income countries such as Ethiopia. Early detection is limited by resource constraints, low screening uptake, and a lack of predictive tools tailored to local healthcare systems. This study presents an interpretable ensemble machine learning framework for predicting CVD risk via structured electronic medical record (EMR) data from public hospitals in Addis Ababa. We trained an XGBoost classifier on 20,960 anonymized records containing demographic, clinical, and physiological attributes. Preprocessing involves handling missing values, outlier capping, one-hot encoding, rare-category grouping, and dimensionality reduction. SHapley additive explanations (SHAPs) were used for feature attribution, and a large language model (Gemini) was used to translate SHAP outputs into plain-language narratives to enhance interpretability. The model achieved an accuracy of 0.99, with strong precision (0.99), recall (0.98), and F1-scores across both classes. SHAP analysis identified general_plan, history of present illness (HPI), musculoskeletal system (MSS) and diagnosis as key predictors. The integration of SHAP and LLMs provided transparent, clinician-friendly insights into model outputs, supporting adoption in resource-limited settings. This study demonstrates that combining ensemble learning with explainability techniques can yield highly accurate and interpretable CVD prediction models, offering potential for integration into clinical decision-support systems in Ethiopia.

创建时间：

2026-02-09

5,000+

优质数据集

54 个

任务类型

进入经典数据集