SEP-PRISM Data: A multi-source dataset for solar energetic particle forecasting
收藏资源简介:
A curated, multi-source dataset for 24-hour-ahead forecasting of operational solar energetic particle (SEP) events. Operational SEP events are defined using the NOAA operational threshold: GOES integral proton flux in the >10 MeV channel reaching or exceeding 10 pfu. The dataset integrates five solar-driver families: Active-region magnetic-field parameters, represented by the unified SMHARP archive Solar flare records Coronal mass ejection (CME) records, represented by the unified CDAWDONKI archive GOES soft X-ray flux in the XRSB 0.1–0.8 nm channel GOES integral proton flux in the >10 MeV channel Ground-truth SEP labels are derived from the CLEAR Benchmark / FetchSEP catalogue. The data products are harmonized into fixed 24-hour historical predictor windows paired with subsequent 24-hour outcome windows, enabling supervised machine learning, benchmark comparison, feature analysis, and operationally motivated SEP forecasting research. Coverage Analyzed rolling-window data: 3 February 1986 to 10 September 2025, UTC Primary prediction horizon: 24 hours Primary event definition: >10 MeV GOES integral proton flux ≥ 10 pfu Primary data product: Analyzed Data/rolling_combinded_seq_24hours.csv Companion publication Yu et al., A multi-source dataset for solar energetic particle forecasting. Preprint: arXiv:2607.16160 Highlights Item Detail Primary ML table Analyzed Data/rolling_combinded_seq_24hours.csv — 14,464 rows × 274 columns (~33 MB) Operational SEP positives 650 / 14,464 rows (4.5%) with Future_OSEP_label = 1 General SEP positives 2,122 / 14,464 rows (14.7%) with Future_GSEP_label = 1 Predictor families Full and flare-matched SMHARP, flare records, CDAWDONKI CME records, CDAW CME records, GOES >10 MeV proton flux, and GOES XRSB flux Extended archives SMHARP extends SHARP-like magnetic predictors to 1996 using aligned SMARP observations; CDAWDONKI extends DONKI-like CME predictors to 1996 using aligned CDAW catalogue information Reproducibility Full Python acquisition + R preprocessing/aggregation pipeline included Repository structure SEP-Prediction-Database/ ├── Raw Data/ # Local workspace for third-party source-data retrieval │ └── README.md # Retrieval instructions; no downloaded raw files included │ ├── Processed Data/ # Harmonized and source-derived data products │ ├── SMHARP.csv # Unified SHARP/SMARP magnetic table (~5.31M rows; ~1.9 GB) │ ├── Flare.csv # Processed flare-event table │ ├── CDAWDONKI_CME.csv # Unified DONKI/CDAW CME table │ ├── CDAW_CME.csv # Processed CDAW CME table │ ├── GOESHAPI_ProtonFlux.csv # Merged GOES >10 MeV proton-flux series │ ├── Merged_XRays.csv # Merged GOES 0.1–0.8 nm X-ray-flux series │ └── All_features.RData # R workspace used by aggregation scripts │ ├── Analyzed Data/ # Model-ready rolling-window data products │ ├── rolling_combinded_seq_24hours.csv # Daily non-overlapping benchmark table (~33 MB) │ └── rolling_combinded_seq_1hours.csv # 1-hour-step overlapping table (~803 MB) │ └── Code/ ├── fetch_solar_data.ipynb # Python workflow for retrieving third-party source records ├── All_features_RData.R # Builds the R workspace from processed data products ├── Data_Aggregation.R # Builds rolling-window machine-learning tables ├── Dataset_Preprocessing/ # Source-specific preprocessing, fusion, and feature extraction └── Plot/ # Publication-figure generation scripts ├── dataspan.R ├── sharp_smarp_plot.R └── Figures/ Raw-data retrieval The public Zenodo archive does not redistribute downloaded third-party raw input files. The Raw Data/ directory is retained as an empty local working directory for the data-retrieval workflow. To reproduce the workflow: Download the CLEAR SEP Benchmark catalogue from the public FetchSEP release and save it in Raw Data/ using the filename expected by the scripts. Run Code/fetch_solar_data.ipynb to retrieve the remaining source records directly from their original providers, including HEK, JSOC, DONKI, CDAW, NOAA SWPC, and NASA ISWA HAPI. The retrieval notebook writes downloaded files into Raw Data/. Run the scripts in Code/Dataset_Preprocessing/, followed by Code/All_features_RData.R and Code/Data_Aggregation.R. Recommended use For standard benchmark experiments, begin with: Analyzed Data/rolling_combinded_seq_24hours.csv Each row contains a 24-hour historical predictor window and targets for the following 24-hour prediction window. The primary binary classification target is: Future_OSEP_label The auxiliary target: Future_GSEP_label supports broader SEP-event modeling, class-imbalance strategies, and multi-task learning. The 1-hour-step analyzed file contains overlapping historical windows; account for temporal dependence and overlap when constructing training, validation, and test splits.



