遇见数据集

Online particle-filter OSSE diagnostics and experiment code for AMOC reconstruction

收藏
Zenodo2026-07-31 更新2026-08-13 收录
官方服务:

资源简介:

Changes in v1.5. The nonlinear observation-operator benchmark (Fig. 7 and Table S3 of the manuscript) was re-run symmetrically. In the v1.4 sweep the particle filter received noisy observations while the linear regression received clean ones, so the two estimators were not compared on equal terms. This version adds lorenz/nl_operator_sweep_symmetric.csv and scripts/nl_demo_symmetric.py, which score both estimators on identical inputs under two symmetric configurations, together with lorenz/nl_operator_sweep_fair.csv and scripts/nl_demo_fair.py. The superseded v1.4 sweep is retained and marked by lorenz/SUPERSEDED_v1.4_nl_operator_note.txt. All other contents are unchanged from v1.4.Data and analysis/experiment code underlying Fallah et al., 'pypfda v1.0: an open-source, model-agnostic particle-filter engine for online sequential paleo data assimilation in Earth system models' (submitted to Geoscientific Model Development, 2026). Contains the perfect-model OSSE diagnostics (truth, FREE and DA ensemble AMOC and subpolar-gyre SST series for the 100-year campaign and the 300-year run; effective-sample-size, importance-weight and full genealogy histories; the 59-site pseudo-observation network), the CLIMBER-X second-core reconstruction, the eta x ESS sensitivity sweep, and the scripts that regenerate every figure and table, including the CM2Mc-BLING SLURM orchestrator (run_online_da.py) and the cost/weight module. The assimilation engine itself is the separate pypfda software record (doi:10.5281/zenodo.21281526). Raw coupled-model output (~16 TB restarts and the 3.25 GB CLIMBER-X truth history) is impractical to archive publicly and is available from the authors on request; see AVAILABILITY.md. [v1.1, 2026-07-10] Added the fair online-vs-offline AMOC benchmark: the exact noisy 59-site pseudo-observations assimilated by the 300-yr run, and scripts reproducing (A) parity between the particle filter and the best linear map from those observations at the assimilation times, and (B) online DA exceeding an offline linear reconstruction interpolated to annual resolution between the sparse observations; plus the nonlinear-observation-operator Lorenz-63 demonstration (particle filter r=0.82 vs linear regression 0.40 under a saturating operator). [v1.2, 2026-07-10] Added the two-panel online-vs-offline figure and the refined analysis: online DA reconstructs the between-observation AMOC variability the sparse 5-yr proxies cannot resolve (ensemble-mean r=0.53 vs 0.28 for an offline linear reconstruction; a perfect-model upper bound), plus the 100-member skill spread; scripts/online_vs_offline.py regenerates the figure and all numbers. [v1.3, 2026-07-28] Recomputed the significance family of the 100-year campaign with null hypotheses that have adequate support, and added the OSSE run configurations. The previous block-permutation null admitted only nine distinct realisations at n=100 with a 10-year block, so counted p-values below 1/10 were unattainable from it; it is retained only as a labelled sensitivity arm, being anti-conservative for a series with a decadal spectral peak. It is replaced by two structure-preserving nulls: an exhaustive circular shift of the truth through all n-1 non-trivial lags, preserving its autocorrelation exactly and enumerated rather than sampled (headline arm p=0.050, k=4 of 99), and a Fourier phase randomisation preserving the full power spectrum (p=0.053, k=5296 of 100000); each surrogate is scored against both the assimilated and the free-running ensemble, so the difference statistic stays paired, and confidence intervals come from a moving-block bootstrap (20000 resamples, 10-year blocks). The effective sample size of the 100-year records is corrected to the two-series Bretherton et al. (1999) estimator (N_eff about 21, critical correlation 0.44); the previously reported value of about 10 came from the single-series form. Under a Benjamini-Hochberg correction across the eight scored campaign arms no arm survives at q<0.05 (smallest q=0.14, shared by the three diverse-initial-condition five-year arms, which are also the only arms whose bootstrap intervals exclude zero): the 100-year campaign ranks the configurations while the inferential claim rests on the independent 300-year integration, whose numbers are unchanged. New files: significance/family_amoc_audited.npz (truth plus all eight campaign arms and both free-running ensembles under the audited alignment convention, with provenance JSON); significance/family_significance_v2.csv and FAMILY_SIGNIFICANCE_V2.md, the file of record for every campaign statistic (correlations, skill differences, confidence intervals, effective sample sizes, per-null p-values with exceedance counts and support sizes, adjusted q-values, mean-square skill scores); scripts/family_significance_v2.py, which regenerates both deterministically from the deposited array and reproduces the deposited CSV byte for byte; scripts/extract_family_amoc_audited.py, documenting the alignment and de-bundling conventions and carrying the calibration gate (r=-0.143, +0.418, +0.494); and run_configs/ with the CM2Mc-BLING input.nml and field_table (background diffusivity 5e-6 m2/s, generic_bling active) and the CLIMBER-X control.nml, closing a gap where the manuscript's Code and Data Availability section stated these were archived here. [v1.4, 2026-07-30] Corrected and completed the CLIMBER-X second-core data, and added model-version provenance for the primary core. In v1.3 the climberx/ directory was internally inconsistent: truth_amoc_climberx.csv held the AMOC truth of a superseded 100-year CLIMBER-X pilot (mean 19.820 Sv, sd 1.130 Sv) while da_result.json in the same directory came from the 500-year production run reported in the manuscript (truth mean 19.088 Sv, sd 1.912 Sv), and no FREE ensemble was deposited, so the manuscript's cross-core skill gain could not be recomputed from the deposit (see climberx/SUPERSEDED_v1.3_note.txt). From this version every file in climberx/ belongs to the 500-year production run, assimilated in 100 cycles of 5 model years, with the deposited AMOC value for each cycle being the cycle-end annual value (model years 5 to 500); the 100-year pilot is not part of the manuscript and is not deposited. Corrected: climberx/truth_amoc_climberx.csv, now the 500-year production-run truth. Added: climberx/free_amoc_members.csv and climberx/da_amoc_members.csv, the full 100-member FREE and DA AMOC ensembles (100 cycles x 100 members); climberx/amoc_ensmean_series.csv, the truth, FREE ensemble mean and DA ensemble mean in one file, the three series the cross-core skill gain is computed from; climberx/free_result.json, the FREE run's driver record (effective sample size, resampling flags, ensemble AMOC), counterpart of the existing da_result.json; and scripts/verify_climberx.py, which recomputes the cross-core skill gain from climberx/*.csv alone and asserts the manuscript's values (raw delta r=+1.39, detrended delta r=+0.98; r_FREE=-0.692, r_DA=+0.701 raw), wired into regenerate_all.sh. Also added run_configs/cm2mc/logfile_T13_DA_5yr_member001.out, the unedited FMS run log of one ensemble member of the headline experiment (T13, diverse initial conditions, five-year cycles), which records the model's own version identifier (MOM_COMMIT_HASH=e664daa1fff332e81b6adf044f50bdc20880f7dd) together with the full namelist as the executable read it, so the source revision behind the reported results is verifiable from this record; the CM2Mc-BLING (Potsdam Earth Model) source itself is GPL-2 code maintained at PIK and is not redistributed here. No CLIMBER-X subpolar-gyre SST series is deposited: per-member SST fields were not accumulated over the 500-year integration (only the AMOC index was) and no CLIMBER-X SST result is reported in the manuscript; subpolar-gyre SST series for the primary core (CM2Mc-BLING) are in amoc_300yr/.

提供机构:
Zenodo
创建时间:
2026-07-31
二维码
社区交流群
二维码
科研交流群
商业服务