遇见数据集

U.S. Life Expectancy and Abridged Period Life Tables by County, State, and Nation, Five-Year Windows 1999-2003 to 2020-2024

收藏
Zenodo2026-08-16 更新2026-08-20 收录
官方服务:

资源简介:

Complete abridged period life tables and life expectancy at birth for every U.S. county, every state, and the nation, as 22 overlapping five-year windows from 1999–2003 through 2020–2024. Tables are for the total population — both sexes combined, all races, all causes — with no demographic or cause-of-death breakdown. Life expectancy at the county level has been estimated before, but the published series stop in 2019 and rest on restricted death records and small-area models that no outside user can reproduce. This series runs through 2024, covering the pandemic and the years after it, and is built entirely from aggregate counts anyone can pull from CDC WONDER, using Chiang's abridged life-table method rather than a statistical model. Files life-expectancy-summary.csv — 70,363 rows, one per area-window. Life expectancy at birth, 65 and 85; standard errors and 95% confidence intervals on e0; lifespan variation measures (e†, SD0, SD10, Keyfitz entropy H); years of potential life lost before 75; and, for every row, the share of deaths directly observed rather than filled. life-tables.csv.gz — 1,336,897 rows, one per area × window × age band. The full abridged life table: n, mx, ax, qx, lx, dx, Lx, Tx, ex, plus an Arriaga decomposition giving each age band's contribution to the gap between that area's life expectancy and the nation's, and two cell-level provenance flags. DATA-DICTIONARY.md — every column defined, with the provenance vocabulary, the fill distribution, and the NCHS validation results. An interactive map, per-county trajectories and per-window tables are at https://demographyinfo.org/life-expectancy. Methods Deaths and population denominators are from CDC WONDER. Life tables are computed by the author using Chiang's abridged life-table method with author modifications; they are not official NCHS or CDC statistics. Five-year pooling is used because single-year county death counts are too sparse to support a stable life table. The life table is fully reproducible from what is published here: Chiang's method needs only mx, ax and the interval width. Limitations — windows overlap and are not independent Consecutive windows share four of their five years, so the series is smoothed by construction and adjacent values are strongly correlated. It is suitable for describing level and broad direction; it should not be treated as 22 independent observations in trend tests or change-point models. Windows from 2016–2020 onward mix pandemic and pre-pandemic mortality, which dampens the 2020–2021 shock relative to single-year estimates. How much is directly measured NCHS suppresses death counts of 1–9, so smaller counties have gaps in their age-specific rates. Suppressed deaths are allocated, withheld population is recovered by subtraction where the arithmetic permits it exactly, and remaining gaps are borrowed from the state mortality schedule. Every cell carries flags recording which: across 1,336,897 cells, 67.9% of death inputs are observed, 21.6% allocated into suppressed cells, and 10.6% borrowed from the state. Every area-window carries Pct_Deaths_Observed — the share of its deaths taken directly from CDC rather than filled — alongside a standard error and 95% confidence interval on e0, so precision can be judged directly rather than inferred from a label. Median values by fill level, for counties: Minimal 99.8% observed and e0 to ±0.31 years (22,652 county-windows); Moderate 85.7% observed, ±0.70 years (28,318); Extensive 62.1% observed, ±1.17 years (15,890); Modeled 85+ 78.1% observed, ±0.58 years (356). Suppressed cells hold 1–9 deaths each by definition, so the fill carries little leverage on e0 outside infancy. The exception is the Capped group (2,003 county-windows), where the median resembles Extensive but the 90th percentile reaches 100% of deaths allocated and ±3.43 years. Those should be read as indicative only, as should any area with many bands borrowed from its state schedule — its life expectancy partly reflects the state's mortality rather than its own. Known issue in version 1.0 — exclude Capped rows 112 of the 70,363 area-windows (0.16%) were computed without a population denominator: 67 in Alaska, 24 in Connecticut and 21 in Virginia, all arising from county boundary and reporting changes. Deaths were present but the matching population was not, and the resulting life expectancies — between 80.1 and 84.7 years — are artifacts rather than estimates. State and national figures are unaffected. All 112 carry Fill_Level = Capped and Data_Quality = Capped. Use Capped as the exclusion criterion, not Pct_Deaths_Observed. That column reads 100% on these rows, because it measures only the share of deaths taken directly from CDC and says nothing about the denominator — so on precisely these rows it indicates high quality where there is none. This is a limitation of the column, and it is stated here rather than left to be discovered. The most conspicuous case is Connecticut, where all eight counties return an identical 84.5 to 84.7 for the windows from 2018–2022 onward, against a Connecticut state value of 79.8, following the 2022 replacement of counties by planning regions in the population estimates. Any county ranking that does not exclude Capped rows will place these eight counties at the top of it. A future version will carry an explicit denominator flag and withhold life expectancy where the denominator is absent. Validation against NCHS State estimates were compared with NCHS, U.S. State Life Tables, 2019 (National Vital Statistics Reports 70-18, Table A), across all 51 states and the District of Columbia. Agreement is close: Pearson r of 0.983 to 0.990 and Spearman ρ of 0.969 to 0.987 depending on the window compared, with per-state mean absolute error of about 0.23–0.26 years once the level offset from five-year pooling is removed. Cross-state dispersion is comparable (SD 1.74–1.91 here against 1.67 in NCHS). One caveat on absolute level. The pre-pandemic 2015–2019 window gives a U.S. e0 of 79.14 against 78.8 published by NCHS for 2019, which is higher than pooling alone comfortably explains. The open interval is the likely source: ax at 85+ is set to 1/mx, standard Chiang and internally consistent, but it assumes a constant hazard above age 85 and tends to overstate survival at the oldest ages. Rankings, gaps and trends are well validated; absolute levels may sit a few tenths of a year high relative to published NCHS figures. What is published, and what is not The age-specific tables here are rates, not counts. Every cell gives mx — deaths per person-year of exposure — together with the life-table quantities derived from it and flags recording whether that cell was observed, allocated or borrowed. The underlying death and population counts are not included. The dimensions are the same as the source extract; the content is not. The distinction is deliberate. A county × age × window table of death counts for 1999–2024 would be a copy of the CDC WONDER extract redistributed outside the system that governs its use. A table of computed rates is analytic output derived from it, of the kind published in any mortality paper. Chiang's method needs only mx, ax and the interval width, so withholding the counts costs nothing in reproducibility; they are needed only to audit the fill arithmetic, and the per-area suppression counts in the summary file preserve the shape of that. Stated plainly, because it is the obvious question: anyone combining these rates with publicly available population denominators could approximately recover death counts for the cells flagged observed. Those are precisely the cells CDC publishes openly — a cell is flagged observed only if it held at least ten deaths and was therefore never suppressed. No suppressed value is recoverable. Cells CDC withheld are flagged allocated or state_rate and carry a fill rather than the true count, so reconstructing them recovers this dataset's estimate, not NCHS's protected figure. This is a terms-of-access decision about bulk redistribution, not a confidentiality one. The full extract remains reproducible from the query parameters documented at demographyinfo.org. Data use — conditions on this dataset The source data were obtained from CDC WONDER under NCHS data use restrictions, which are accepted at the point of download and are not carried by a file. Because that acceptance step cannot travel with a Zenodo deposit, the substance of it is imposed here instead. By downloading these files you accept the following, in addition to the CC BY-NC-SA 4.0 licence. 1. Use these data for statistical reporting and analysis only. 2. Make no attempt to learn the identity of any person or establishment included in the underlying vital records, and do not link these data with any other dataset for the purpose of identifying an individual. 3. If you inadvertently discover the identity of an individual, make no disclosure or use of it, and report the circumstance to NCHS. 4. Observe the same restrictions in anything you redistribute, and pass this notice on with it. These conditions follow the NCHS restrictions published at https://wonder.cdc.gov/datause.html, which govern the source data and should be read in full. Nothing here narrows them. On the nine-or-fewer rule. The NCHS restrictions direct users to avoid publishing statistics based on nine or fewer cases. NCHS suppresses death counts of 1–9 in WONDER, so those cells arrive as gaps rather than numbers. Where a rate is published for such a cell, it is an allocated fill — not a rate computed from the withheld count, which was never available — and every one is flagged allocated or state_rate in Deaths_source. The flag itself conveys only what WONDER already states publicly, that the cell held fewer than ten deaths. Users should not treat any flagged cell as a measured rate, or report a figure resting on one without carrying the flag alongside it. Both files, and the per-area Pct_Deaths_Observed and Suppressed_Cells columns, exist so that this is visible in every row rather than buried in a caveat. Data available on request The intermediate tables — raw and enhanced age-specific counts, per-window build outputs, and the extraction parameters — are available on request, subject to CDC WONDER's terms of use, to anyone who asks through demographyinfo.org. Requests are answered. The full extract is also reproducible from the documented query parameters. Source data Centers for Disease Control and Prevention, National Center for Health Statistics. National Vital Statistics System, Mortality data on CDC WONDER Online Database. Accessed at https://wonder.cdc.gov/ in July 2026; final extraction 31 July 2026. Data were retrieved from more than one WONDER mortality dataset across the 1999–2024 span; dataset identifiers, groupings, and year ranges are documented at demographyinfo.org. Because no race grouping is applied, total deaths and total population are identical across the bridged-race and single-race series, so the change of dataset has no effect on these tables. Suggested citation Lahey, T., & DemographyInfo.org. U.S. Life Expectancy and Abridged Period Life Tables by County, State, and Nation, Five-Year Windows 1999–2003 to 2020–2024. Version 1.0, 31 July 2026. https://demographyinfo.org/life-expectancy. Accessed [date].

提供机构:
DemographyInfo.org
创建时间:
2026-08-16
二维码
社区交流群
二维码
科研交流群
商业服务