Processed Data and Trained Models Supporting "Machine Learning Reconstruction of Glider Dissolved Oxygen: Accuracy Across Missions, Depths, and Years"
收藏资源简介:
This Zenodo record contains the processed data and trained models supporting the manuscript “Machine Learning Reconstruction of Glider Dissolved Oxygen: Accuracy Across Missions, Depths, and Years.” The files support the model evaluations, dissolved oxygen reconstructions, interannual analyses, and numerical results reported in the article and its Supporting Information. They were prepared specifically for this study and do not replace the original observational archives. The data include 41 northern Gulf of Mexico glider missions during 2013–2025, including M51 in 2025, together with ship-towed observations from the Mechanisms Controlling Hypoxia (MCH) project and shipboard profiles from the Southeast Area Monitoring and Assessment Program (SEAMAP). Twelve glider missions supplied usable dissolved oxygen (DO) measurements. The other 29 missions provided predictor observations for DO reconstruction. Processed variables include DO, temperature, salinity, pressure, in situ seawater density, optical backscatter, fluorescent dissolved organic matter (fDOM), chlorophyll-a, position, sampling time, bottom depth, and an upper-water-column stability predictor. Screening flags, missing-value indicators, and sample identifiers document observation quality, predictor availability, and membership in training and validation datasets. Measured DO and reconstructed values are stored separately. The uploaded files are: 01_processed_data.zip: Quality-screened observations, model input tables, sample identifiers, prepared sequence inputs, scaling parameters, and missing-predictor filling parameters. Mission inventories, glider tracks, MCH transects, and SEAMAP station information describe observational coverage. Filling and scaling parameters were estimated from the training observations for each experiment. 02_final_trained_models.zip: Four final trained models: a multilayer perceptron (MLP), an adapted Self-Attention-based Imputation for Time Series network (SAITS), random forest (RF), and histogram gradient boosting (HGB). These models correspond to the first of five repeated fits, which supplies the reconstruction fields displayed in the article. The MLP uses a softplus output to constrain DO concentrations to be nonnegative and was trained with randomly hidden optical inputs to improve reconstruction when those inputs are missing. The archive includes joblib and PyTorch model files, filling and scaling parameters, training settings, and software-environment information. Intermediate validation-model checkpoints are not included. 03_predictions_and_reconstruction.zip: Validation predictions and one-minute glider reconstruction tables for all five repeated fits. The final-third test used the first two-thirds of observations, ordered by time, within each mission, section, or annual survey for training and its final third for validation. The mission-exclusion test withheld entire glider missions from training. Sample identifiers allow predictions from the two tests to be compared on the same observations. Reconstruction tables retain measured DO where available and provide estimates from each model where DO is missing, including M51 in 2025. 04_article_statistics.zip: Validation statistics by observation source, mission, region, and pressure; monthly climatologies; annual DO and temperature anomalies; errors in reconstructed annual DO anomalies; interannual variability and trend estimates; model-specific lower-layer definitions; missing-optical-input sensitivity tests; and summaries across the five repeated fits. README.txt: File organization, principal variables, units, model-loading requirements, original data sources, and interpretation limits. DO concentrations in the processed model tables are stored in micromoles per kilogram (µmol kg⁻¹). Concentrations and errors reported in millilitres per litre (mL L⁻¹) were calculated sample by sample using in situ seawater density. Salinity is unitless, pressure is in dbar, and temperature is in degrees Celsius. Reconstruction tables retain screening flags; the article uses rows marked input_qc_pass = True. Annual anomalies were calculated by subtracting the corresponding monthly climatological means and accounting for differences in mission, spatial, and monthly sampling. Each annual series was then centered by subtracting its mean across years. Missing months were not filled. Variation across repeated fits describes sensitivity to model fitting and is not a complete estimate of reconstruction uncertainty. This record does not include publication figures, analysis source code, development logs, or complete copies of the original observational archives. Code will be made available separately. Loading the trained models requires the dependencies and custom model definitions specified in the README. The data and models should be used with the screening criteria, validation design, and limitations described in the associated article. Reconstructed DO is not an independent observation, and agreement among models does not establish reconstruction accuracy.



