遇见数据集

Development and validation of a prediction model for long-term cognitive frailty risk in stroke patients based on CHARLS data

收藏
Zenodo2026-01-04 更新2026-05-26 收录
官方服务:

资源简介:

I. Basic Information of the Dataset Dataset name:Data for Development and validation of a prediction model for long-term cognitive frailty risk in stroke patients based on CHARLS data.csv Data source: Based on the China Health and Retirement Longitudinal Survey (CHARLS) data from 2018 to 2020 Sample size: It includes a sample of 2,325 stroke survivors Data format: CSV (Comma-Separated Values) format, structured table data Core application: To construct and validate a risk prediction model for long-term cognitive frailty in stroke patients Ii. Background of Data Collection The data is derived from the survey waves of CHARLS in 2018 and 2020, using a multi-stage scale-proportional sampling (PPS) method, covering 150 counties in 28 provinces across the country Inclusion criteria: Community-dwelling elderly people aged 60 and above who have been diagnosed with stroke by a doctor Exclusion criteria: No history of stroke, Alzheimer's disease/dementia, severe mental illness, severe aphasia/dysarthria, severe hearing impairment, and advanced severe systemic diseases (such as advanced heart failure, end-stage liver disease) Iii. Variable Description (A total of 23 Variables) 1. Target variable CF (Cognitive Frailty) : Cognitive frailty state (1 = presence of cognitive frailty, 0 = no cognitive frailty), the prevalence of cognitive frailty in the sample was 29.59% (688/2325). Variable sort Classification description CF(Cognitive Frailty) 0 no 1 yes Age 1 60~ 2 70~ 3 ≥80 Education 1 below primary school 2 primary school 3 junior high school 4 high school and above Gender 0 female 1 male Marital status 0 unmarried/divorced/widowed 1 be married Permanent address 0 urban 1 rural Self-reported health status 0 less satisfied/average 1 satisfied IADL 0 normal 1 impaired Number of chronic diseases 1 0~ 2 2~ 3 ≥4 Drinking 0 no 1 yes Smoking 0 yes 1 no BMI 0 ≤24 1 ≥24 Nutritional status 0 malnutrition 1 well-nourished Life satisfaction 0 less satisfied/average 1 satisfied Chronic pain 0 no 1 yes Sleep quality 0 good 1 poor Social event 0 yes 1 no Intellectual activity 0 yes 1 no Live alone 0 no 1 yes Self-rated life satisfaction 0 less satisfied/average 1 satisfied History of Falls 0 no 1 yes Exercise 0 no 1 yes Depressed 0 no 1 yes Iv. Data Quality and Processing Missing value handling: Missing values ≤30% are filled with R language mice packages (continuous variables are matched with predicted means, and categorical variables are matched with logistic regression). Outlier handling: Evaluated and adjusted based on the clinically acceptable range Data preprocessing: Standardization of continuous variables and individual hot encoding of categorical variables Data segmentation: Stratified sampling was adopted to divide the data into the training set (70%, 1627 cases) and the test set (30%, 698 cases). V. Data Association and Application Key correlations: Age, educational level, nutritional status, physical exercise, and IADL are core predictors of cognitive frailty (screened by LASSO regression) Main application: It is used for training and validating 8 types of machine learning prediction models (logistic regression, decision tree, XGBoost, etc.), among which the XGBoost model performs the best (AUC=0.810) Scalability: Supports the analysis of risk factors for cognitive frailty, optimization of predictive models, and research on clinical intervention strategies

提供机构:
Zenodo
创建时间:
2026-01-04
二维码
社区交流群
二维码
科研交流群
商业服务