遇见数据集

An Explainable Neuro-Fuzzy Model for Hypertension Prediction

收藏
Mendeley Data2026-09-08 收录
官方服务:

资源简介:

This dataset consists of 1249 clinical records created for research on explainable machine learning and neuro-fuzzy modelling for hypertension prediction. The data set combines demographic, anthropometric, clinical, lifestyle and family history data that are relevant to hypertension risk classification. The data set includes 14 variables: Age, Gender, Height (cm), Weight (kg), Body Mass Index (BMI), Systolic Blood Pressure, Diastolic Blood Pressure, Cholesterol (mg/dL), Blood Glucose (mg/dL), Smoking, Alcohol Intake, Physical Activity, Family History, and Hypertension Status. The Hypertension_Status variable is the target variable for hypertension classification. The data set was prepared for the study “An Explainable Neuro-Fuzzy Model for Hypertension Prediction” at Osun State University, Osogbo, Nigeria. The accompanying research methodology is based on the selection of clinical variables for model development, which are Age, BMI, Systolic Blood Pressure, Diastolic Blood Pressure, Cholesterol Level, and Blood Glucose Level. The proposed modelling approach is a combination of fuzzy reasoning and machine learning to deal with the nonlinear relationships and uncertainty in clinical prediction. The research methodology involves data preprocessing, categorical encoding, feature selection, feature scaling, partitioning of the dataset, development of a neuro-fuzzy model, baseline machine learning modelling, performance evaluation, and explainability analysis using SHAP. The proposed neuro-fuzzy system is based on Sugeno-type fuzzy inference system with linguistic variables, membership functions and expert defined IF–THEN rules. The data set is designed to facilitate academic research, experimentation, benchmarking, machine learning development, explainable artificial intelligence research, and investigation of intelligent approaches for hypertension risk prediction. It is not intended to be used as a clinically validated diagnostic tool and predictions made by models trained on this data should not be used as a substitute for clinical assessment or clinical decision making.

本数据集包含1249条临床记录,旨在开展高血压预测相关的可解释机器学习与神经模糊建模研究。该数据集整合了与高血压风险分类相关的人口统计学、人体测量学、临床、生活方式及家族病史数据。 本数据集共包含14项变量:年龄(Age)、性别(Gender)、身高(单位:厘米,Height (cm))、体重(单位:千克,Weight (kg))、体重指数(Body Mass Index, BMI)、收缩压(Systolic Blood Pressure)、舒张压(Diastolic Blood Pressure)、胆固醇(单位:mg/dL,Cholesterol (mg/dL))、血糖(单位:mg/dL,Blood Glucose (mg/dL))、吸烟(Smoking)、饮酒(Alcohol Intake)、体力活动(Physical Activity)、家族病史(Family History)以及高血压状态(Hypertension Status)。其中,高血压状态为高血压分类任务的目标变量。 本数据集由尼日利亚奥绍博奥孙州立大学开展的《用于高血压预测的可解释神经模糊模型(An Explainable Neuro-Fuzzy Model for Hypertension Prediction)》研究项目整理制备。配套的研究方法围绕模型开发所需的临床变量筛选展开,筛选出的变量包括年龄、BMI、收缩压、舒张压、胆固醇水平及血糖水平。所提出的建模方法结合了模糊推理与机器学习技术,用以处理临床预测任务中的非线性关系与不确定性。 该研究方法涵盖数据预处理、类别编码、特征选择、特征缩放、数据集划分、神经模糊模型开发、基准机器学习建模、性能评估以及基于SHAP的可解释性分析。所提出的神经模糊系统基于杉野型模糊推理系统(Sugeno-type fuzzy inference system),包含语言变量、隶属度函数及专家定义的IF–THEN规则。 本数据集旨在推动高血压风险预测相关的学术研究、实验验证、基准测试、机器学习模型开发、可解释人工智能研究以及智能方法探索。本数据集并非经过临床验证的诊断工具,基于本数据集训练得到的模型所生成的预测结果,不应替代临床评估或临床决策。

创建时间:
2026-08-18
二维码
社区交流群
二维码
科研交流群
商业服务