遇见数据集

<p>Descriptive characteristics of dataset.</p>

收藏
NIAID Data Ecosystem2026-05-10 收录
官方服务:

资源简介:

Objective Child stunting continues to pose a substantial global health challenge, requiring multifaceted strategies that combine conventional epidemiological approaches with advanced analytic methods. The aim of this study was to determine the most effective machine learning model for predicting stunting based on water, sanitation, and hygiene behaviors and infrastructure, with the goal of identifying high-risk children who would benefit most from targeted interventions. Methods This study was a secondary analysis of data from a matched cohort study assessing the effectiveness of combined on-premise piped water and improved sanitation for improved health outcomes in rural Odisha, India. Data for the parent study were collected from 2,398 households with a child under five years of age across 90 villages, and complete data were available for 1,196 children. Feature engineering techniques were employed to identify the most relevant predictors and utilized structural equation modeling, forward selection, backward elimination, and least absolute shrinkage and selection operator techniques. Five machine learning algorithms commonly used for binary classification tasks were compared: logistic regression, classification tree, support vector machine, neural network, and extreme gradient boosting. Results Among 1,196 children analyzed, the extreme gradient boosting model with forward selection feature engineering best predicted stunting based on water, sanitation, and hygiene (WaSH) factors. It correctly identified 81% of stunted children and 92% of non-stunted children, with an overall accuracy of 88%. The model’s area under the receiver operating characteristic curve (AUROC) was 0.959 (95% CI: 0.949–0.968), indicating that WaSH factors strongly predict child stunting when analyzed using this advanced machine learning technique. Four WaSH factors were identified as having the strongest power to predict stunting in our sample: improved sanitation coverage, presence of a handwashing station, piped water coverage, and availability of preferred drinking water source. Conclusions The results demonstrate the efficacy of machine learning algorithms, especially extreme gradient boosting to potentially inform targeted WaSH interventions for reducing childhood stunting in resource-limited settings. However, these findings require external validation in other populations, and the complete-case analysis approach (excluding 35% of children with missing data) may limit generalizability to settings with less systematic data collection.

### 研究目的 儿童生长迟缓仍是一项严峻的全球公共卫生挑战,亟需结合传统流行病学方法与先进分析技术的多维度干预策略。本研究旨在基于水、环境卫生与个人卫生(Water, Sanitation and Hygiene,以下简称WaSH)相关行为与基础设施条件,筛选出预测儿童生长迟缓效果最优的机器学习模型,以识别最能从针对性干预中获益的高风险儿童群体。 ### 研究方法 本研究为一项匹配队列研究的二次数据分析,该队列研究旨在评估印度奥里萨邦农村地区联合使用户内管道供水与改良卫生设施对改善健康结局的效果。母研究共从90个村庄的2398户有5岁以下儿童的家庭中收集数据,最终获得1196名儿童的完整数据集。本研究采用特征工程技术筛选最具相关性的预测因子,并运用结构方程模型、向前逐步选择法、向后剔除法以及最小绝对收缩和选择算子(Least Absolute Shrinkage and Selection Operator,LASSO)技术。本研究对比了5种常用于二分类任务的机器学习算法:逻辑回归、分类树、支持向量机、神经网络以及极端梯度提升(Extreme Gradient Boosting,XGBoost)。 ### 研究结果 在本次分析的1196名儿童中,结合向前逐步选择特征工程的极端梯度提升模型对基于WaSH因素的儿童生长迟缓预测效果最优。该模型可正确识别81%的生长迟缓儿童与92%的非生长迟缓儿童,整体准确率达88%。其受试者工作特征曲线下面积(Area Under the Receiver Operating Characteristic Curve,AUROC)为0.959(95%置信区间:0.949–0.968),表明采用该先进机器学习技术分析时,WaSH因素可有效预测儿童生长迟缓。本研究共筛选出4项对样本内儿童生长迟缓预测能力最强的WaSH因素:改良卫生设施覆盖率、洗手台配备情况、管道供水覆盖率以及优质饮用水源可及性。 ### 研究结论 研究结果证实了机器学习算法的有效性,尤其是极端梯度提升模型,可为资源受限环境下开展针对性WaSH干预以降低儿童生长迟缓提供决策参考。但上述研究结论需在其他人群中进行外部验证,且本研究采用的全案例分析方法(排除了35%存在数据缺失的儿童)可能限制了研究结果推广至数据收集系统性较弱的场景。

创建时间:
2026-03-05
二维码
社区交流群
二维码
科研交流群
商业服务