Sample survey dataset for Machine learning models in Diabetes prediction using socio economic factors
收藏资源简介:
The primary dataset used in this study is collected through a survey conducted in three districts of Assam, namely Nagaon, Lakhimpur, and Kamrup Rural. The survey aimed to capture socio-demographic, lifestyle, and health-related variables relevant to non-communicable disease (NCD) risk factors, with a particular focus on diabetes. The final analytic sample consisted of 791 individuals. Prior to model development, the dataset underwent a systematic pre-processing procedure to improve data quality, eliminate information leakage, and ensure compatibility with the machine learning (ML) algorithms. Administrative identifiers, including participant name, household code, individual code, village identifiers, timestamps, and other non-informative variables, were removed because they did not contribute to prediction and could potentially compromise participant anonymity. A complete-case analysis approach was adopted, whereby observations containing missing values in any of the variables used for modelling were excluded from the analysis. This resulted in a final analytical sample of 791 individuals, ensuring that all ML models were trained using a consistent dataset without the need for statistical imputation.



