遇见数据集

Reproduction for figures in the Influpaint paper "Generative diffusion models for spatiotemporal influenza forecasting"

收藏
Zenodo2026-09-11 更新2026-10-01 收录
官方服务:

资源简介:

Data and figures for the Influpaint paper Zenodo DOI: 10.5281/zenodo.22699980. Paths below are relative to the extracted archive unless a command specifies the research repository root. The Zenodo reproducibility archive contains the pretrained model, forecast outputs, observations, and analysis data used to generate the paper's figures and supplementary figures. It also contains the regenerated figures. Forecasts and model weights are provided as archived outputs; plotting them does not require retraining the model or generating new forecasts. This archive is a companion to the Influpaint GitHub repository, which contains the code. From the root of a cloned Influpaint repository, using the Influpaint Python environment, run: python -m paper_figures.final_figures --data-root "influpaint-paper/influpaint_paper_reproduction_data" This single command generates the eight paper/supplementary data plots, the Figure 1 correlation summary JSON, and the FluSight leaderboard summary CSV and comparison plot in regenerated_paper_figures/ inside the archive. It uses random seed 0 by default; use --seed to select another seed or --output-dir to select another output directory. The input tables, model weights, and forecast arrays are archived data, not outputs regenerated by this plotting command. Source dataset datasets/sources/all_datasets.parquet contains 2,535,111 standardized source rows from FluView, FluSurv, flepiMoP round 1, and Flu Scenario Modeling Hub rounds 4–5. Columns identify the source family, model/scenario, season, sample, location, season week, date, and value. This is a copy of the locally available gathered-source table, last modified November 7, 2025. It is not established as the original source-table realization used to create the July 17 training datasets. Training datasets datasets/training/ contains the exact July 17, 2025 NetCDF files selected by the paper configuration: File Mixture recipe Saved frames TS_100S_2025-07-17.nc Surveillance only 520 TS_70S30M_2025-07-17.nc 70% surveillance / 30% modeling 3,332 TS_30S70M_2025-07-17.nc 30% surveillance / 70% modeling; selected paper dataset 3,223 TS_100M_2025-07-17.nc Modeling only 1,240 Dimensions are (sample, feature, season_week, place), with shape (N, 1, 64, 64). The 53 season-week positions and 51 state/DC positions are padded to 64 × 64. Files retain main_origins and mix_cfg metadata. Mixture percentages describe the recipes; realized frame counts differ from the requested mixing size. datasets/provenance.json records source repository paths, byte sizes, and SHA-256 hashes for the source Parquet, four training NetCDFs, and model_candidate_evaluation/scoringutils_scores.csv. Model checkpoint i868::m_U500cRx1224::ds_30S70M::tr_Sqrt::ri_No::3000.pth is the pretrained checkpoint for the final model formulation used in the paper's retrospective forecasting and generation analyses: Setting Value Model identifier i868 Full formulation i868::m_U500cRx1224::ds_30S70M::tr_Sqrt::ri_No Diffusion process 500 steps with a cosine noise schedule U-Net ResNet blocks with channel multipliers (1, 2, 2, 4) Training-data mixture 30% surveillance and 70% simulated trajectories Observation transform Square-root scaling Additional training enrichment None Training duration 3,000 epochs Forecast conditioning CoPaint, configuration celebahq_noTTJ5 The file contains the trained network weights, optimizer state, epoch setting, and loss-type metadata. The network architecture and preprocessing are defined in the Influpaint source code. The original checkpoint filename was i868::m_U500cRx1224::ds_30S70M::tr_Sqrt::ri_No::3000.pth, from the paper-2025-07-22_training experiment. Retrospective forecasts forecasts/retrospective/ contains forecasts from the paper's selected model for 29 reference dates across the 2023–2024 and 2024–2025 influenza seasons. Each date has its own directory containing: A CSV file: the four-week forecast distribution in FluSight format, summarized using 23 quantiles. These values produce the short-horizon forecast fans in Figure 2. fluforecasts_ti.npy: the full-season ensemble of 512 trajectories, including the observed-period portion and the future portion. These samples produce the full-season forecasts in Figure 3. ti denotes the inverse-transformed output: values are on the hospitalization-count scale. The CSV columns are: Column Meaning reference_date Reference date identifying the forecast target Forecast target, wk inc flu hosp horizon The four target weeks, encoded as 0, 1, 2, and 3 in these files target_end_date End date of the week being predicted location FluSight location code output_type quantile output_type_id Quantile probability, including 0.5 for the median value Predicted weekly incident influenza hospitalizations at that quantile Use target_end_date to identify the forecasted week: in these CSVs, horizon 0 has a target end date equal to the reference date. Quantiles summarize the distribution separately for each target; the NPY files preserve complete sampled trajectories. Unconditional generated seasons forecasts/unconditional/ contains inverse_transformed_samples_i868::m_U500cRx1224::ds_30S70M::tr_Sqrt::ri_No.npy. These are 512 complete synthetic influenza seasons generated by the pretrained model without conditioning on an observed season. They are the source of the trajectories, uncertainty bands, and cross-state correlation analysis in Figure 1. Masking and reconstruction experiments forecasts/masks/ contains the model outputs for the six masking experiments shown in Figure 4. All six use the 2023–2024 season (season2023 in the directory names). Directory Reconstruction task missing_half_subpop_season2023 Reconstruct trajectories for approximately half the locations while observing the others missing_nc_season2023 Reconstruct North Carolina with that state's observations withheld missing_il_season2023 Reconstruct Illinois with that state's observations withheld missing_midseason_biggap_season2023 Reconstruct a missing interval within the season missing_past_season2023 Reconstruct the early season using observations from later in the season missing_checkerboard_4x4_season2023 Reconstruct missing blocks in a checkerboard pattern of four weeks by four locations Each directory contains: fluforecasts_ti.npy: 512 reconstructed trajectory samples on the hospitalization-count scale. mask.npy: the conditioning mask. 1 identifies observed entries supplied to the model; 0 identifies hidden entries to reconstruct. ground_truth.npy: the untransformed observation array saved with the experiment, against which reconstructions can be compared. Missing observations may be represented by NaN. NPY array coordinates The retrospective, unconditional, and reconstructed sample arrays all have shape (512, 1, 64, 64), with axes sample, channel, week, location. There is one hospitalization channel. The mask and saved ground-truth arrays have shape (1, 64, 64), with axes channel, week, location. The model pads the week and location dimensions to 64. The first 51 spatial positions represent the 50 states and Washington, DC; the remaining positions are padding. Their order follows observations/influpaint_locations.csv after excluding the national US entry and territories, as done by SeasonAxis.for_flusight(remove_us=True, remove_territories=True). Use the Influpaint SeasonAxis calendar, the season identifier, and observation dates to interpret the time dimension. The season calendar starts in August. Padded positions are not additional weeks or locations to include in epidemiological summaries. National trajectories in Figure 3 are obtained by summing the 51 state/DC trajectories within each sample. Operational forecasts and comparison forecasts forecasts/operational/ contains selected CSVs copied from the official FluSight forecast-hub repositories for each season: UNC_IDD-InfluPaint/: forecasts actually submitted during the season. These produce the operational forecast fans in Supplementary Figure 4. FluSight-ensemble/: the FluSight ensemble forecasts used as comparisons in Figure 2 and Supplementary Figure 4. These are historical submissions. They are distinct from the retrospective forecasts in forecasts/retrospective/, and should not all be attributed to the single archived i868::m_U500cRx1224::ds_30S70M::tr_Sqrt::ri_No::3000.pth checkpoint. Observations and location metadata File Contents and source observations/2023-2024/target-hospital-admissions.csv Observed weekly influenza hospital admissions, copied from the 2023–2024 FluSight hub's target data; used for the corresponding forecast and reconstruction comparisons observations/2024-2025/target-hospital-admissions.csv Observed weekly influenza hospital admissions, copied from the 2024–2025 FluSight hub's target data; used for the corresponding forecast comparisons observations/nhsn_flusight_past.csv Historical NHSN/FluSight hospitalization series, copied from influpaint/data/nhsn_flusight_past.csv; includes season and season-week fields and supplies the historical curves and observed correlations in Figure 1 observations/influpaint_locations.csv Location codes, names, abbreviations, population, and geographic identifiers, copied from the Influpaint package; used to associate array positions with states These are snapshots of the local observation files used by the plotting code. A file may include dates outside the season named by its directory; the plots select the appropriate dates and locations. Location codes should be read as strings to preserve leading zeros. Training and model-comparison results analysis/ contains the training-loss tables needed for Supplementary Figures 1–3. Model-comparison scores and rankings are in model_candidate_evaluation/. File Contents and use mlflow_loss_timeseries.csv Training loss by logged step, with scenario/run identifiers and timestamps; used for the training-loss curves mlflow_losses.csv Per-run training summaries, including final loss, average loss over the last 100 steps, and training metadata; used to relate training loss to forecast skill The two loss tables were exported from the training experiments. These tables include candidate model formulations as well as the selected i868 formulation. WIS denotes weighted interval score; lower values indicate better forecast accuracy. Model candidate evaluation model_candidate_evaluation/ is a complete copy of the research repository's folder of the same name. It contains: Path within the folder Contents scoringutils_scores.csv Saved per-forecast scores for all candidates and FluSight comparison models leaderboards/leaderboard_full.csv WIS sums, mean relative WIS, and ranks for 36 InfluPaint formulations in each season and combined; used by the supplementary paper figures simple_plots/2023-2024/, simple_plots/2024-2025/, simple_plots/Combined/ Diagnostic heatmaps, performance comparisons, time series, and state-level WIS components simple_plots/interactive_model_selection.html Interactive comparison of candidate scores simple_plots/paper_model_analysis.csv Selected-model and FluSight comparison summary flusight_dropbox_analysis.csv Summary of the archived operational FluSight leaderboards plot_evaluation_results.log Console output from the diagnostic plotting run To regenerate the candidate evaluation directly inside this archive, run from the cloned research repository root with the Influpaint Python environment: MPLBACKEND=Agg python -u -m evaluation.plot_evaluation_results \ --csv-path influpaint-paper/influpaint_paper_reproduction_data/model_candidate_evaluation/scoringutils_scores.csv \ --save-dir influpaint-paper/influpaint_paper_reproduction_data/model_candidate_evaluation/simple_plots \ --leaderboard-dir influpaint-paper/influpaint_paper_reproduction_data/model_candidate_evaluation/leaderboards \ > influpaint-paper/influpaint_paper_reproduction_data/model_candidate_evaluation/plot_evaluation_results.log 2>&1 This regenerates 45 diagnostic PNGs, the leaderboard, interactive comparison, paper summary, and log from the saved scores. The script excludes i808, UGuelph-CompositeCurve, and CADPH-FluCAT_Ensemble, and applies its per-season completeness rule. Time-series plots show the top 3 or top 10 models per group; full comparisons include all eligible models. The repository supplies the forecast-job manifest and location metadata used by the plotting code. Image rendering can vary slightly with library versions. The nine flusight_*ranked.png and flusight_relative_wis_wis_dual.png files are retained FluSight-only plot snapshots from the earlier evaluation. The default command above does not regenerate those snapshots. The current script's --group-filter flusight --annotate-ranks options generate the dual-metric chart; use a separate --save-dir for that run to avoid replacing the full candidate comparison plots. The two standalone ranked chart types are not generated by the current script. To regenerate the operational leaderboard summary (also producing a WIS comparison plot): python -m paper_figures.analyze_flusight_dropbox_tables \ --data-root influpaint-paper/influpaint_paper_reproduction_data \ --output-dir influpaint-paper/influpaint_paper_reproduction_data/model_candidate_evaluation To generate the local research output instead, copy the archived model_candidate_evaluation/ folder to the repository root and run MPLBACKEND=Agg python -m evaluation.plot_evaluation_results; its defaults use that local folder. The scores themselves were generated by joining forecast quantiles to observations with python -m evaluation.prepare_dataset_for_scoringutils, then running: Rscript evaluation/score_with_scoringutils.R \ model_candidate_evaluation/combined_forecast_truth_data.csv \ model_candidate_evaluation/scoringutils_scores.csv Recomputing these scores requires the full set of candidate and comparison forecast CSVs and configured source paths. The archive contains only the selected model's retrospective forecasts, so it supports reproducing the complete comparison from saved scores, but not rescoring every candidate from raw forecasts. Paper figures regenerated_paper_figures/ contains the eight data plots included in the paper and supplement: File Figure 868_figure1_unconditional_correlation.png Figure 1: unconditional generation and correlations 868_figure2_csv_forecasts_two_seasons.png Figure 2: four-week retrospective forecasts 868_figure3_npy_forecasts_two_seasons.png Figure 3: full-season forecast trajectories 868_figure4_mask_experiments.png Figure 4: masking and reconstruction experiments sup_forest_effect.png Supplementary Figure 1: model-formulation comparisons sup_training_loss.png Supplementary Figure 2: training-loss curves sup_lossVSwis.png Supplementary Figure 3: training loss versus forecasting performance 868_Relaizedforecast.png Supplementary Figure 4: actual operational FluSight submissions; Relaized is a retained filename typo Figure 1 correlation summary The same command also writes regenerated_paper_figures/868_figure1_correlation_summary.json. This is computed from the exact correlation values plotted in Figure 1b. Each of the following entries contains count, mean, and median: JSON entry Meaning compute_random_correlation Null correlations from 100 draws of a generated season, independently permuting each state's weeks compute_weekly_incidence_correlation Pairwise state correlations pooled over the 512 unconditional generated seasons compute_observed_correlation Pairwise state correlations pooled over the observed 2023–2024 and 2024–2025 seasons The calculation uses the 51 state/DC locations, excluding spatial padding, and Pearson correlations over weekly incidence. Pairs with fewer than three jointly observed weeks or effectively constant trajectories are excluded. The observed input is observations/nhsn_flusight_past.csv; generated and null correlations use forecasts/unconditional/. The generated and observed means reproduce the manuscript's rounded values of 0.646 and 0.833. Null statistics depend on the random seed. FluSight leaderboard tables and analysis analysis/FlusightScores/ contains the source leaderboard snapshots used for the operational-performance rankings in the manuscript: File Source 2022-2023_table_flusight_mathispaperSI.txt 2022–2023 FluSight evaluation table from the Mathis paper's supplementary information 2023-2024_table_flusight_dropbox.txt 2023–2024 FluSight Dropbox leaderboard snapshot 2024-2025_table_flusight_dropbox.txt 2024–2025 FluSight Dropbox leaderboard snapshot The code is paper_figures/analyze_flusight_dropbox_tables.py in the Influpaint repository. The reproduction command above runs it automatically and writes these additional files in regenerated_paper_figures/: File Contents flusight_dropbox_analysis.csv InfluPaint's absolute WIS, relative WIS, MAE, ranks, coverage, and comparison-model counts for all three seasons 2024-2025_wis_pairgrid.png Absolute and relative WIS comparison across eligible 2024–2025 models All seasons exclude model names starting with FluSight (case-insensitive). The 2023–2024 and 2024–2025 analyses also require at least 70% of forecasts submitted; the 2022–2023 analysis uses the supplied evaluation table without that additional threshold. Coverage columns retain the source tables' labels (50% Coverage (%) and 95% Coverage (%)). These outputs summarize the archived leaderboard values; they do not rescore raw forecasts.

提供机构:
Zenodo
创建时间:
2026-09-11
二维码
社区交流群
二维码
科研交流群
商业服务