遇见数据集

U.S. Harmonized Vital Statistics (HVS) Microdata: Natality, Linked Birth–Infant Death, Fetal Death, and Matched Multiples

收藏
Zenodo2026-05-28 更新2026-05-29 收录
官方服务:

资源简介:

The U.S. Harmonized Vital Statistics (HVS) resource integrates four National Center for Health Statistics (NCHS) public-use vital-statistics microdata products into stable Apache Parquet column schemas that span multiple Standard Certificate revisions (1989 and 2003) and NCHS layout reformats. Each product carries one harmonized schema across every year it covers, removing the per-year fixed-width layout, field-name, and code-value drift that has historically forced researchers to restrict analyses to single-revision windows. Four products (this deposit): - Natality — 1968–2024 (57 years), 201,161,456 records, 84 columns (71 harmonized + 13 derived). Validated 183/183 Births: Final Data NVSR targets byte-exact for 1990–2024. - Linked birth–infant death — 1983–2023 (38 cohort years; permanent 1992–1994 NCHS-linkage gap), 149,386,620 records, 97 columns. - Fetal death — 1982–2024 (43 years), 2,427,233 records, 89 columns. - Matched multiples — twins/triplets/quadruplets with linked infant and fetal deaths across three NCHS publication windows (1995–1997, 1995–2000, 2016–2020), 1,665,568 records, 24 columns. Files: seven primary harmonized/derived Apache Parquet files (~8.3 GB), plus SHA256SUMS.txt, README.md, LICENSE, and CITATION.cff. Per-year raw parquets are reproducible from the GitHub pipelines and public NCHS source zips. Validation: every shipped per-year figure is reconciled against the corresponding National Vital Statistics Reports (NVSR) cell under each product’s documented canonical analytic filter; per-target PASS/FAIL tables ship with the resource on GitHub. Reproducibility: open-source pipelines (MIT; see Related works) re-derive every parquet bit-for-bit from public NCHS source zips. SHA-256 checksums are in each product’s PROVENANCE.md and in SHA256SUMS.txt. Licensing: harmonized data CC BY 4.0; source code MIT. Underlying NCHS source data are U.S. Government works (17 U.S.C. § 105). No NCHS Research Data Center access is required for this public-use envelope. Geography note: NCHS suppresses sub-national geography in these public-use files; the resource supports national and demographic-stratum analyses, not state-level analyses. Supersedes: prior single-product Zenodo deposits 10.5281/zenodo.19363074 (natality + linked) and 10.5281/zenodo.20031571 (fetal death 1992–2022); those records remain immutable.

美国统一生命统计(Harmonized Vital Statistics, HVS)资源将美国国家卫生统计中心(National Center for Health Statistics, NCHS)的四款公共使用生命统计微数据产品整合为稳定的Apache Parquet列式架构,覆盖多版标准出生证明修订版(1989年与2003年)及NCHS布局改版。各产品在其覆盖的所有年份中采用统一的标准化架构,消除了历史上迫使研究者仅能在单一修订版时间窗口内开展分析的逐年固定宽度布局、字段名称与编码值的变动偏差。 本资源包含四款产品: - 活产数据集:覆盖1968–2024年(共57年),包含201,161,456条记录,共84个字段(其中71个为统一字段,13个为衍生字段)。经183/183项出生数据验证通过,1990–2024年的数据与《国家生命统计报告》(National Vital Statistics Reports, NVSR)的最终数据实现字节级完全匹配。 - 出生-婴儿死亡关联数据集:覆盖1983–2023年(共38个队列年;存在1992–1994年永久性NCHS关联缺口),包含149,386,620条记录,共97个字段。 - 胎儿死亡数据集:覆盖1982–2024年(共43年),包含2,427,233条记录,共89个字段。 - 多胎匹配数据集:涵盖双胞胎/三胞胎/四胞胎,关联婴儿与胎儿死亡数据,覆盖三个NCHS发布窗口(1995–1997、1995–2000、2016–2020),包含1,665,568条记录,共24个字段。 文件列表:7个核心统一/衍生Apache Parquet文件(总大小约8.3 GB),外加SHA256SUMS.txt、README.md、LICENSE及CITATION.cff。逐年原始Parquet文件可通过GitHub流水线及公开NCHS源压缩包复现生成。 验证机制:所有发布的逐年数据均按照各产品文档中规定的标准分析过滤器,与对应《国家生命统计报告》(NVSR)中的单元格数据进行比对校验;各验证目标的通过/失败表随本资源一同发布于GitHub。 可复现性:采用开源流水线(MIT协议;详见相关研究),可从公开NCHS源压缩包逐比特复现生成所有Parquet文件。SHA-256校验和存储于各产品的PROVENANCE.md文件及SHA256SUMS.txt中。 许可协议:统一后的数据采用CC BY 4.0协议;源代码采用MIT协议。底层NCHS源数据属于美国政府作品(依据美国版权法第17编第105条)。本公共使用数据集无需申请NCHS研究数据中心访问权限。 地理说明:NCHS在本系列公共使用文件中隐去了次国家级地理信息;本资源仅支持全国层面及人口分层分析,不支持州级层面分析。 本资源替代了此前单产品的Zenodo发布包:10.5281/zenodo.19363074(活产+关联数据集)及10.5281/zenodo.20031571(1992–2022年胎儿死亡数据集);上述原有记录仍保持不可修改。

提供机构:
Zenodo
创建时间:
2026-05-28
二维码
社区交流群
二维码
科研交流群
商业服务