Reproducible Workflow for "Calibration-Target Variability Shapes Machine-Learning Reconstructions of Seawater δ18O From Paired Coral Sr/Ca–δ18O Records
收藏资源简介:
This collection contains the reproducible Python workflow, processed data, model outputs, figures, and repository tables supporting the manuscript “Calibration-Target Variability Shapes Machine-Learning Reconstructions of Seawater δ18O From Paired Coral Sr/Ca–δ18O Records.” It includes analyses based on CoralHydro2k, PAGES seawater δ18O observations, EN4 sea-surface salinity, grouped machine learning, and pseudoproxy experiments. Table R1. Complete record inventory, site-group assignments, sampling-frequency screening, and PAGES target-feasibility diagnostics. Complete inventory and screening diagnostics for the 53 CoralHydro2k records extracted using the exact LiPD variable names SrCa and d18O and paired by calendar month. The table reports record metadata, temporal coverage, the number of paired months and mean proxy values, the maximum absolute difference between the mean decimal dates of Sr/Ca and coral δ18O values assigned to the same calendar month, and sampling-frequency metrics calculated from gaps between consecutive paired months. coverage_fraction is the number of paired months divided by the inclusive record span. Records are classified as monthly dominant when the median gap is ≤1 month and at least 80% of gaps are exactly 1 month; bimonthly dominant when the median gap is ≤2 months and at least 80% of gaps are ≤2 months; annual or lower resolution when the median gap is ≥10 months; and irregular otherwise. monthly_primary_eligible identifies monthly-dominant records before the common-reference, target-availability, and final machine-learning eligibility filters were applied. Target feasibility summarizes nearby surface or unknown-depth PAGES seawater δ18O support: good, moderate, and weak require at least five observations within ≤250, ≤500, and ≤1,000 km, respectively; sparse indicates one to four usable observations, and unusable indicates none. These feasibility classes describe observational support and do not by themselves determine final machine-learning inclusion. Sr/Ca is reported in mmol mol−1, isotope values in ‰, time gaps in months, and distances in km. Table R2. Global and basin δ18Osw–SSS regression diagnostics. Ordinary least-squares regressions relating PAGES seawater δ18O (δ18Osw) to salinity, used as global and basin fallback calibrations when constructing the EN4/SSS-derived calibration pseudo-target. Regressions have the form δ18Osw = intercept + slope × SSS and use observations with 20 ≤ SSS < 45 psu and water depth ≤50 m or unknown. Global and basin fits require at least 30 observations, an SSS range of at least 0.2 psu, a finite coefficient, and a positive slope; qc_pass records whether these safeguards were met. Reported diagnostics include coefficient standard errors, residual standard deviation, Pearson correlation (r), coefficient of determination (r2), RMSE, observed SSS range, and the mean and standard deviation of observed δ18Osw. Slope is in ‰ psu−1; intercept, residual standard deviation, RMSE, and δ18Osw statistics are in ‰. target_window is a pipeline-provenance label for the primary EN4 extraction-window run; the PAGES-derived regression coefficients were reused unchanged in the extended-window sensitivity analysis. Table R3. Selected δ18Osw–SSS regressions and record-level target diagnostics for the primary benchmark. Selected regression and target-variability diagnostics for the 17 coral records retained in the primary benchmark constructed from the 1980–2016 EN4 extraction window, with eligible coral–target observations spanning 1980–2013. For each record, an accepted local regression within 1,000 km was selected when it contained at least 20 observations, spanned at least 0.2 psu, and had a finite positive slope; otherwise, an accepted basin regression and then the accepted global regression served as fallbacks. The table reports the selected regression scope, coefficients and standard errors, residual standard deviation, sample size, r, r2, RMSE, and SSS range, together with the target-source identifier. Record-level means and population standard deviations (ddof = 0) of the EN4/SSS-derived calibration pseudo-target, the unscaled traditional estimate, the target-minus-traditional residual, and EN4 SSS are calculated over the months entering the machine-learning-ready dataset. regression_rmse_to_calibration_pseudo_target_sd_ratio is the selected-regression RMSE divided by the temporal standard deviation of the record-level calibration pseudo-target; it measures calibration scatter relative to target variability and is not a machine-learning prediction-error metric. Table R4. Controlled pseudoproxy results across calibration-target variability levels. Ensemble results for six target-retention values (λ = 0.20, 0.30, 0.40, 0.50, 0.75, and 1.00) in the controlled pseudoproxy experiment. Each row summarizes 30 independently generated networks, each containing 15 synthetic sites and 300 monthly observations per site, evaluated using leave-one-site-out predictions from the ridge residual-correction model. target_retention is the nominal multiplier λ applied to the true δ18Osw signal before independent target noise was added and the target was centered within each site; consequently, the realized calibration-target-to-true-signal SD ratio need not equal λ exactly. Columns ending in _mean are ensemble means, and _q05 and _q95 are the 5th and 95th percentiles across the 30 networks. Performance is evaluated separately against the available calibration target and the known true signal. ml_gain_vs_traditional_rmse_against_true_signal is defined as traditional-estimate RMSE minus machine-learning RMSE, so positive values favor machine learning. RMSE values are in ‰; correlations and SD ratios are dimensionless. Within each replicate, the same proxy network and true signal were used across λ values, whereas target-noise realizations were generated separately. Table R5. Primary-benchmark performance of the designated model and transparent comparison methods. Pooled leave-one-site-group-out performance for the primary benchmark constructed from the 1980–2016 EN4 extraction window, comprising 4,292 record-month observations from 17 records in 15 validation site groups and spanning an eligible coral–target overlap of 1980–2013. The a priori designated model is the traditional-inclusive elastic-net residual-correction model. Comparators are the unscaled fixed-coefficient traditional estimate, zero anomaly, the outer-training-group mean, the outer-training-group monthly climatology, and training-only linear shrinkage of the traditional estimate; all training-dependent quantities were estimated without the held-out site group. Metrics are calculated from pooled out-of-fold predictions. bias is mean(prediction − calibration pseudo-target), and prediction_to_target_sd_ratio is SD(prediction)/SD(target), using population standard deviations (ddof = 0). rank is ordered by increasing RMSE. Differences relative to the unscaled traditional estimate are calculated as method minus traditional: negative ΔRMSE and negative percentage RMSE change indicate lower error, whereas positive Δr indicates higher correlation. Correlation is undefined for the zero-anomaly baseline because its prediction is constant. Table R6. Performance of the a priori designated model and 11 exploratory model–feature combinations. Pooled leave-one-site-group-out performance of six model families—mean-residual dummy, ridge, elastic net, random forest, extra trees, and gradient boosting—under two predictor modes in the primary benchmark constructed from the 1980–2016 EN4 extraction window, with eligible coral–target overlap spanning 1980–2013. The core mode uses 20 proxy, sampling, lag, seasonality, and location predictors; the traditional-inclusive mode adds the contemporaneous unscaled traditional δ18Osw anomaly as a 21st predictor. The traditional-inclusive elastic-net model was designated a priori as the primary residual-correction model, whereas the remaining 11 combinations were exploratory sensitivity analyses. is_designated_model identifies the designated combination. RMSE, MAE, r, and prediction-to-target SD ratio are calculated from pooled out-of-fold predictions across the 15 held-out site groups, and rank is ordered by increasing RMSE. The unscaled traditional-estimate metrics are repeated for reference. Identical mean-residual-dummy results under the two feature modes are expected because that model does not use predictor values. validation_group_column_key records site_group as the outer validation unit. Table R7. Pooled performance and variability diagnostics for the primary and extended EN4 extraction windows. Sensitivity comparison of the 1980–2016 and 1950–2016 EN4 extraction windows used to construct the calibration pseudo-target. The complete target-construction, feature-generation, grouped model-tuning, leave-one-site-group-out prediction, and baseline workflow was rerun for each window. The same PAGES δ18Osw–SSS calibration hierarchy, 1980–1989 common anomaly reference, 17 coral records, and 15 validation site groups were retained. The corresponding eligible coral–target overlaps were 1980–2013 and 1950–2013, respectively, and the number of eligible record-month observations increased from 4,292 to 8,404. Metrics are calculated from pooled out-of-fold predictions using population standard deviations (ddof = 0). target_to_traditional_sd_ratio is SD(calibration pseudo-target)/SD(unscaled traditional estimate), ml_to_target_sd_ratio is SD(designated ML prediction)/SD(calibration pseudo-target), and corr_predicted_correction_vs_traditional_estimate quantifies residual cancellation by correlating the predicted correction with the traditional estimate. Table R8. Paired site-group inference for the primary and extended EN4 extraction windows. Paired comparisons of the designated machine-learning model with training-only linear shrinkage and zero anomaly across the 15 held-out site groups in each EN4 extraction-window experiment. mean_delta_rmse is defined as site-group RMSE for machine learning minus site-group RMSE for the stated baseline; negative values favor machine learning, and ml_better_n counts site groups with negative differences. ci_low and ci_high are the 2.5th and 97.5th percentiles of 10,000 nonparametric bootstrap estimates of the mean difference obtained by resampling site groups with replacement. wilcoxon_p is the two-sided paired Wilcoxon signed-rank p value calculated from the same site-group RMSE differences. Site groups, rather than individual records or months, are the inferential units.



