Harmonised Spanish National Health Survey (ENSE/EESE/ESdE) 2006-2023
收藏资源简介:
This record provides an individual-level harmonised dataset that pools seven waves of the Spanish national health interview surveys conducted between 2006 and 2023 by the National Statistics Institute (INE) and the Ministry of Health: the Encuesta Nacional de Salud de España (ENSE 2006, 2011–2012 and 2017), the Encuesta Europea de Salud en España (EESE 2009, 2014 and 2020; the Spanish component of the European Health Interview Survey, EHIS) and the Encuesta de Salud de España (ESdE 2023). It contains 160,937 respondents aged 16 years or older and 121 variables: four identification and weighting variables and 117 harmonised variables in 11 domains. The harmonised variables cover sociodemographics, proxy interviews, self-perceived health and activity limitation (GALI), 18 chronic conditions (each as ever, last 12 months and doctor-diagnosed), activities of daily living (65+), health care use, medicines, prevention, self-reported anthropometry, lifestyles and socioeconomic position. Household income is calibrated to the INE Household Budget Survey (EPF) and comes with a within-wave relative-position (ridit) variable for inequality analyses. Two weights are provided: the INE adult weight (FACTOR), for population estimates within a wave, and a weight normalised within each wave (FACTORUA), for pooled analyses. Comparability is rated in two stages:- Ex ante harmonisation level, from the comparison of questionnaires: 40 variables are identical (level 1) and 77 similar (level 2) across waves. Constructs with substantive differences were not included.- Ex post caution level, specific to each wave, derived from the data: 60 variables have caution level 1 (standard use), 37 level 2 (caution) and 20 level 3 (high caution in the indicated waves). The rating combines (i) an internal assessment of 90 indicators (215 indicator–wave combinations) that tests whether deviations from sex- and age-standardised trends are consistent across sex–age groups and are associated with a documented question change or with the survey mode (telephone interviewing in 2020; web then face-to-face in 2023), and (ii) an external validation against official statistics (177 indicator–wave comparisons; INE Continuous Population Statistics, Labour Force Survey, Municipal Register, Household Budget Survey and Living Conditions Survey; Eurostat and EU-SILC). Suspect waves are flagged, not removed. Findings of the assessment were used to adjust the harmonisation post hoc (regroupings, corrections, new derived variables, calibration of household income, and exclusion of one construct). Every departure from the literal coding of the INE files is documented in the recoding table and can be switched on or off with a named flag in the code. Note that analyses by household income that include 2020 or 2023 require high caution (caution level 3), because the ranking of households is much less informative in those waves and cannot be corrected with the public files. The record contains:- data_harmonized: the harmonised dataset in SPSS (.sav, with variable and value labels), R (.rds) and CSV (UTF-8, numeric codes) formats;- documentation: master tables (literal wording of every question in every wave) and recoding tables (source variables, recoding rules and missing-value codes by wave), in Spanish and English; a weighted descriptive report by wave, sex and age; a comparability report with the harmonisation and caution level of every variable; and an external validation report, each with machine-readable CSV versions;- R: the full pipeline that rebuilds the dataset from the INE and Ministry of Health public-use microdata (wave-specific scripts, pooling script and run-all script), [and the scripts of the internal assessment and external validation];- a README file with the list of input files and versions. Variable names and value labels are in Spanish, as in the questionnaires; the English meaning of each variable and category is given in the English master and recoding tables. Missing values are coded as system-missing. Two variable names contain the character ñ, so the code must be run in a UTF-8 session. The public-use files do not include the primary sampling units, so standard errors based on the weights alone ignore the sampling design. The source microdata are not redistributed. They were downloaded on 26 September 2026 (latest versions available at that date) and can be obtained free of charge from INE (www.ine.es) and the Ministry of Health. Producers occasionally revise their files, so re-running the code requires the versions listed in the README. This dataset is a derived work and neither INE nor the Ministry of Health is responsible for it. Please cite this record and the original surveys.



