Municipal dengue risk dataset for Mexico, 2022: incidence, climatological, environmental, sociodemographic, socioeconomic and agricultural covariates
收藏资源简介:
Municipal-level dengue incidence for Mexico in 2022, integrated with climatological, environmental, sociodemographic, socioeconomic and agricultural covariates, together with the ordinal risk variable derived from it. The file covers 1,050 municipalities across all 32 Mexican states, contains 36 columns and has no missing values. The 1,050 municipalities are those with complete 2022 precipitation and temperature records; the remaining municipalities of the country were excluded because these two climatological predictors could not be imputed without introducing uncertainty. They are geographically interspersed with the retained ones rather than forming a distinct region. The dataset provides the reported dengue incidence rate (Rate_2022, cases per 100,000 inhabitants), the municipal centroid coordinates, thirty candidate covariates, and the Dengue Risk Level (DRL), an ordered variable with five levels. Level 0 includes the 580 municipalities with no reported cases; the 470 with positive incidence are divided at the empirical quartiles of the positive distribution, with cut-points of 7.105, 24.15, and 73.00 cases per 100,000, giving four levels of 118, 117, 117, and 118 municipalities. These cut-points are a relative stratification specific to this sample and this year, not officially established epidemiological thresholds, and level 0 indicates that no cases were reported rather than that transmission was absent. Covariates come from public Mexican sources: precipitation and temperature from CONAGUA; sociodemographic, housing and agricultural variables from INEGI, drawing on the 2020 Population and Housing Census and the 2022 agricultural statistics; poverty indicators from CONEVAL. Incidence records derive from Mexico’s General Directorate of Epidemiology through the Supporting Information of Mendoza-Cano et al. (2025). Two composite indices, SEVI and SEAVI, are derived here as the first principal component of groups of standardized variables that are themselves included, so both can be recomputed. An accompanying data dictionary documents the role, units, source, reference year, and observed range of every column. This version supersedes versions 1 and 2, which contained only the fifteen covariates retained by a feature-selection step performed on the complete sample, and which omitted the state and municipality identifiers. The analysis code that uses this dataset is available at https://github.com/sausolofer/Ordinal-Regression-Kriging



