遇见数据集

County-Level COVID-19 Cases with Population, Density, Stringency, and Temperature Covariates: A Merged Source Panel for the First US Pandemic Wave (2020)

收藏
Zenodo2026-07-14 更新2026-08-01 收录
官方服务:

资源简介:

Overview. This dataset is a linked county-day panel for the contiguous United States covering 22 January – 15 July 2020, the first wave of the COVID-19 pandemic. Each record corresponds to one county on one calendar date and includes the reported case count for that day, along with time-invariant county attributes (population, land area, density, coordinates, monthly mean temperature) and the prevailing state-level policy stringency. The dataset assembles four public sources into a single analysis-ready table, including the raw and linked variables prior to any modeling. Data sources. Four publicly available sources were linked on county, state, and date. County-level daily reported case counts were obtained from the Johns Hopkins Coronavirus Resource Center. Resident population (2020 estimates base), county land area (square miles), and population density (persons per square mile) were obtained from the US Census Bureau Population Estimates. Daily policy stringency was obtained from the Oxford COVID-19 Government Response Tracker as the weighted-average Stringency Index, resolved at the state level and broadcast to each constituent county. Monthly mean county temperature for January through July 2020 was obtained from the National Center for Environmental Information. Each source retains its original license; users should cite the primary providers in addition to this deposit. Linkage and keying. The sources were joined on a composite county key (county name, state, and country) and, for the time-varying case and stringency series, on calendar date. County identity is preserved in both human-readable form (county, state) and as a combined key, together with latitude and longitude, to support geographic linkage to external datasets. Cleaning and quality control. Index artifacts and duplicate columns introduced during source merging were removed, and duplicate county-day records were collapsed. Alaska and Hawaii were excluded because several county-level fields had a high proportion of missing values, thereby restricting the dataset to the contiguous United States. Case counts are retained as reported by the source; no smoothing, imputation, or outlier adjustment was applied at this stage, so that downstream users can apply their own pre-processing. Variable schema. Each row contains: county and state identifiers, country, latitude and longitude, and a combined location key; the reported daily case count and its calendar date; the 2020 population estimate, county land area (sq mi), and derived population density (persons/sq mi); the state-level weighted-average Stringency Index for that date; and seven columns of monthly mean temperature (January–July). Coverage and structure. The panel spans 22 January – 22 July 2020 at daily resolution across US counties. Time-invariant county attributes (population, land area, density, coordinates, monthly temperatures) are repeated across a county's daily rows by construction; the two genuinely time-varying fields are the daily case count and the daily stringency index.

提供机构:
Zenodo
创建时间:
2026-07-14
二维码
社区交流群
二维码
科研交流群
商业服务