Harmonized U.S. Natality and Linked Infant Death Microdata, 1990-2024
收藏资源简介:
v2.7.0 (this release) combines the v2.6 schema upgrade and v2.7 pre-release audit fixes since the v2.5.0 deposit. A harmonized dataset of 138,819,655 U.S. birth records (1990-2024) and 74,943,824 linked birth-infant death records (2005-2023), derived from NCHS public-use natality files. The pipeline resolves five fixed-width record layouts, two birth-certificate revisions, and dozens of field-name and coding changes into a single stable schema with explicit cross-year comparability documentation. Schema: V2 natality is 71 harmonized + 13 derived = 84 columns. V3 linked is 78 harmonized + 16 derived = 94 columns (V2 plus 7 death-side harmonized + 3 derived). Validation: 183 of 183 V2 NVSR external targets pass; 35 of 35 V3 linked user-guide targets pass; 41 of 41 internal invariants pass on V2 (V3 linked: 38 pass clean, 1 within a documented 2-row exception budget, 3 V2-only invariants skipped per documented file-format differences). What changed since v2.5.0: v2.6 (schema upgrade): Schema grew from 82 to 84 columns (V2) and 92 to 94 columns (V3) with two new variables (father_age_cat_from_rec11 and maternal_race_detail_15cat). Field-position fixes restored about 14 million previously-null cells (2004 ATTEND read from byte 408 instead of 410; 2013 FAGECOMB at bytes 182-183 and RF_CESAR at byte 324). 2016 onward, diabetes and hypertension now derive from RF_PDIAB/RF_GDIAB instead of the stale URF tail block. 2016 onward, the linked-cohort merge uses the composite (CO_SEQNUM, CO_YOD) key per NCHS guide. v2.7 (audit fixes): Documentation drift corrected across paper drafts and schema CSV, including per-era field counts (37/35/36/43/75), invariant counts (41), 1990-2002 smokers-with-unknown-intensity (428,755), and URF/RF fallback year (2016+). Three new V3 caveats added to the schema: maternal_race_bridged4 coverage ends 2019 (NCHS dropped the MBRACE field in 2020+); V3 payment_source_recode and father_education_cat4 are 100% NULL in 2009-2010 because the LinkCO09 and LinkCO10 zips ship with those bytes blank (verified by raw-byte probe). Validator V3 mode-detection hardened to require both infant_death and record_weight; output filenames now carry v2 / v3_linked mode tags. Convenience writer Kleene-trap fixed and PROVENANCE preservation logic added. Harmonizer --years defaults restored to full ranges. Data integrity: parquet row content is byte-identical between v2.5.0 and v2.7.0 for the columns v2.5.0 already had; v2.7.0 adds the v2.6 columns and restored cells. Verify against SHA-256 checksums in PROVENANCE.md. Reading order for new researchers: README.md, then GETTING_STARTED.md, CODEBOOK.md, COMPARABILITY.md, VALIDATION.md, and FAQ.md. The quickstart.ipynb notebook has working examples. REPRODUCING.md explains how to rebuild from raw NCHS source files using the open-source pipeline at https://github.com/yoelplutchok/natality-harmonization Note on file-internal DOI references: Some files inside this deposit (README.md, REPRODUCING.md, FAQ.md, ABOUT_THIS_RELEASE.md, quickstart.ipynb, requirements.txt) contain links written as 10.5281/zenodo.19363075 — that is v2.5.0's version-specific DOI, not the concept DOI. The correct concept DOI for "always the latest version of this dataset" is 10.5281/zenodo.19363074. The GitHub repository (https://github.com/yoelplutchok/natality-harmonization) and the v2.7.0 release notes (https://github.com/yoelplutchok/natality-harmonization/releases/tag/v2.7.0) cite the correct concept DOI. This documentation inconsistency will be corrected in the next release.



