遇见数据集

Interpretable Machine Learning for Daily Precipitation Forecasting: Spatiotemporal Dynamics and Extreme Event Prediction in New Jersey

收藏
Zenodo2025-11-25 更新2026-05-26 收录
官方服务:

资源简介:

This dataset contains the complete, processed data used in the manuscript titled "Interpretable Machine learning for Enhanced Daily Precipitation Forecasting: Spatiotemporal Dynamics and Extreme Event Prediction in New Jersey." The data is provided in a single comma-separated value (CSV) file named Final_ML_Dataset_All_Features.csv. Dataset Characteristics: Temporal Coverage: 2015-01-01 to 2024-12-31 Spatial Coverage: 31 meteorological stations across New Jersey, USA. Total Observations (Rows): 113,243 Total Variables (Columns): 40 (39 predictor features + 1 target variable) Format: Station-wise daily time series. Each row represents a single observation for one station on a specific day. Content:The dataset integrates multiple data sources to provide a rich set of predictors for one-day-ahead daily precipitation forecasting. It includes: Target Variable: Precipitation (daily total in mm) sourced from NASA GPM/IMERG. Meteorological Predictors: Daily variables such as 2m air temperature, dew point temperature, surface pressure, and wind components, sourced from the ERA5-Land reanalysis dataset. Engineered Temporal Features: Lagged variables (e.g., precip_lag1), rolling window statistics (e.g., precip_roll3), and calendar-based features (e.g., day_of_week). Static Environmental Features: Topographic variables (elevation, slope) from the SRTM 30m DEM, and remote sensing indices (NDVI, NDBI) for the years 2015 and 2024 derived from MODIS and Sentinel-2 imagery. Station Metadata: station name, latitude, and longitude. A complete data dictionary detailing all 39 predictor variables is provided in the Appendix of the associated manuscript. This dataset was used to train and evaluate six machine learning models (LR, SVR, RF, ADA, KNN, XGB) and to conduct the interpretability and spatiotemporal analyses presented in the paper. All numerical predictor variables in this dataset are provided in their raw (non-normalized) format. The preprocessing steps, including Z-score normalization, are described in the manuscript's methodology section. This dataset is shared to ensure the reproducibility of the study's findings and to facilitate future research in hydrological forecasting and interpretable machine learning.

提供机构:
Zenodo
创建时间:
2025-11-25
二维码
社区交流群
二维码
科研交流群
商业服务