Seasonal probabilistic forecasting of reservoir storage droughts in Peru: data and reproducible methodology
收藏资源简介:
Reproducibility archive (version 2) for the study Seasonal probabilistic forecasting of reservoir storage droughts in Peru: skill and economic value against the operational precipitation-based warning (manuscript submitted for peer review, 2026). This version (2): doi:10.5281/zenodo.22900614 · All versions: doi:10.5281/zenodo.22082641 Peru's operational drought warning is built on a multi-scale Standardized Precipitation Index (SPI). It precedes almost every reservoir storage drought only because it is active in about 86% of months, so it cannot tell an operator which month requires action. This study forecasts the probability of storage drought one to six months ahead for 20 reservoirs in three hydroclimatic regions, using each reservoir's own storage record, and verifies the forecasts against the operational warning on identical reservoir-months. What changed in version 2. Version 1 verified the forecasts by cross-validation with indices standardised over the full record. Version 2 replaces that design with a progressive validation in which nothing uses future information: for each issue year, the storage anomaly, the precipitation indices, the seasonal-forecast anomaly, the models and the sensitivity thresholds are fitted with data up to 31 December of the previous year and then frozen. All results were recomputed with this stricter design and supersede those of version 1. The Poechos operation counterfactual of version 1 is not carried forward, because it relied on indices standardised over the full record. This archive contains: codigo.zip — the methodological code chain, five Python scripts numbered in execution order: SEAS5 retrieval from the Copernicus Climate Data Store; basin aggregation of the SEAS5 ensemble; past-only standardisation, reconstruction of the operational SPI warning and the M0–M5 forecast models under progressive validation with frozen yearly fits; verification on identical reservoir-months (ROC area, Hanssen–Kuipers score with fixed and past-selected thresholds, Brier and CRPS skill, reliability, cost–loss economic value, year-block bootstrap); and derived products (alert frequency against skill with the arithmetic ceiling KSS ≤ (1−q)/(1−s), skill by domain and stratum, storage memory, sample composition and an illustrative case). No absolute paths need editing; the scripts write to salidas/ and never overwrite the published results. datos.zip — the published results: every forecast issued for 2009–2026 with its predictors, predictive distribution, drought probability, the reconstructed warning and the observed outcome (predicciones_verificadas.parquet), from which any verification metric can be recomputed without refitting; all verification tables; the yearly fitting diagnostics; the derived products; and the basin-aggregated SEAS5 predictor. README.md — documentation in English: contents, script-by-script description, the validation protocol, how to reproduce, external inputs, environment, license and citation. LEEME.md is the Spanish counterpart. LICENSE.txt — Creative Commons Attribution 4.0 International. How to reproduce: (1) unzip codigo.zip and datos.zip at the same package root; (2) place the external inputs listed below under datos_externos/; (3) run scripts 03, 04 and 05 in order (about three minutes on a laptop) and compare salidas/ with datos/. Scripts 01 and 02 are needed only to rebuild the SEAS5 predictor from the original GRIB files. External inputs (not redistributed here): the monthly storage series, basin precipitation, reservoir metadata and contributing basins of the companion storage-drought dataset (doi:10.5281/zenodo.21697481); the ICEN (ENFEN) and ONI (NOAA/CPC) indices from their official sources; and SEAS5 forecasts, retrieved free of charge from the Copernicus Climate Data Store with the included script. Key results supported by this archive: on identical reservoir-months (issue months January 2017 to June 2026; 1,781 reservoir-months at lead 1), the storage forecast reaches ROC areas of 0.94, 0.79 and 0.70 at leads of 1, 3 and 6 months against 0.52–0.47 for the reconstructed operational warning, while alerting in 22–37% of months instead of 85–86%; a warning that is active in 86% of months cannot exceed a Hanssen–Kuipers score of 0.16; the forecast probabilities are reliable and have positive cost–loss value at all 19 ratios up to lead 3; and skill at lead 6 follows reservoir memory, being retained in the long-memory reservoirs of the southern Andes. Environment: Python ≥3.11 with numpy, pandas, scipy, scikit-learn, lightgbm, pyarrow, xarray, cfgrib, geopandas, shapely and cdsapi. License: CC-BY-4.0 (external inputs retain the licenses of their original sources). When using this material, please cite this version (doi:10.5281/zenodo.22900614) or all versions of the deposit (doi:10.5281/zenodo.22082641), together with the associated article.



