GRINS AQCLIM: The GRINS datasets on Air Quality and CLimate for Italy
收藏资源简介:
--> Make sure you are downloading the latest version: v.2.0.0 <-- 1. GRINS_AQCLIM datasets 1.1 GRINS_AQCLIM_points_Italy dataset The GRINS_AQCLIM_points_Italy dataset record contains datasets about the measurements of air pollutant concentrations, recorded by the Italian ground-based monitoring network, along with related climate variables, uncertainties and stations-related information. There are mainly three datasets: GRINS AQCLIM dataset, Station registry information, GRINS AQCLIM imputation uncertainty. This dataset provides daily summary statistics for numerous air pollutants, logged at 744 locations over an eleven-year period, from 2013 to 2023. The pollutants considered include: nitrogen monoxide and dioxide (NO and NO2), particulate matter (PM10 and PM2.5), ozone (O3), ammonia (NH3), carbon monoxide (CO), and sulfur dioxide (SO2). Further details are available on the dedicated GitHub page: https://github.com/GRINS-Spoke0-WP2/AQ-EEA. The dataset are organized into two primary categories: Air Quality (AQ) and Climate (CL). This division is reflected in the column naming convention, where each column name includes a corresponding prefix (e.g., "AQ_" or "CL_"). The dataset provides a unique space-time identification from: "AirQualityStation": The ID code of the monitoring station "time": Represents the recording day. For the Air Quality (AQ) dimension, there are 48 columns formed by combining the "AQ_" prefix with various summary statistics (minimum, first quartile, mean, median, third quartile, and maximum, denoted as "min", "q1", "mean", "med", "q3", "max") for each considered pollutant. As an illustration, the column name "AQ_q1_NO2" signifies the first quartile of the daily nitrogen dioxide distribution recorded by that particular station. Regarding the Climate (CL) dimension, the variables included are detailed in Table 2. Comprehensive information about these climate variables is accessible on the Copernicus Climate Data Store website: https://cds.climate.copernicus.eu/datasets/reanalysis-era5-single-levels?tab=overview. The dataset is available in rda and CSV format. To facilitate optimal management and accommodate the considerable overall data volume, the dataset in CSV format has been segmented into biennial periods. The specific naming helps the identification of the period covered. 1.2 GRINS_AQCLIM_Station_registry_information dataset The dataset contains useful information on the monitoring stations that should be coupled to the overall dataset. We kept this information in a separate file because they are not changing over time, in order to optimize storage. The Station_registry_information dataset contains: "AirQualityStation": The ID code of the monitoring station "Longitude" and "Latitude": Provide the geographical coordinates of the station, using the WGS-84 reference system. "AirQualityStation": The ID code of the monitoring station "Altitude": Indicates the station's altitude. "AirQualityStationType": Describes the predominant type of emission sources in the station's vicinity. "AirQualityStationArea": Specifies the type of area surrounding the station. 1.3 GRINS_AQCLIM_imputation_uncertainty Hourly data were converted to daily resolution. In order to avoid bias we have imputed some missing hourly data using a local level model and the Kamlan smoother. We notice that the days with more than five consecutive missing hourly data are set as missing in the daily dataset. In order to take into account for the uncertainty of the missing-data imputation strategy, we provided the GRINS AQCLIM imputation uncertainty dataset. These estimates are calculated considering both variances and covariances through all the daily summary statistics. In particular, if the minimum or maximum daily values are imputed, the square root of the Kalman smoother conditional variance is reported. For quartiles and median, only conditional variances and covariances of the two selected ordered statistics are considered. For the daily average, all the conditional variances and covariances intra-day are considered. We notice that zero values refer to daily averages without missing values at the hourly level, NaN values refer to missing daily averages, and positive uncertainties are related to daily averages with one or more hourly imputed values. In the GRINS AQCLIM imputation uncertainty dataset are reported only the days when at least one imputation is done, to optimize storage. Columns are names the same as the GRINS AQCLIM dataset with a "sd_" in front of each of them. 2. GRINS_AQ_NO2_LAUs dataset The dataset contains daily concentrations of nitrogen dioxide for Italian municipalities. Predictions are obtained from the implementation of the Fixed Rank Kriging on air quality data from the EEA monitoring stations, along with several external variables as weather and emissions. Output grid (0.05°x0.05°) is then aggregated to municipal level. Period covered is 2013-2023. Area is Italy. PRO_COM codes are referred to 2025 administrative Italian municipalities borders. The GRINS_AQ_NO2_LAUs dataset contains: "PRO_COM": The ID code of the municipality "mean_NO2_FRK": The aggregated prediction obtained from the FRK model "sd_NO2_FRK": The aggregated uncertainty obtained from the FRK model "time": The day considered 3. GRINS_CL_LAUs dataset The dataset contains daily climate variables for Italian municipalities. Data are obtained from ERA5-Land and ERA5 datasets. Due to a better resolution, ERA5-Land is preferred to ERA5 when available. Grids (0.1°x0.1°) are aggregated to municipal level. Period covered is 2013-2023. Area is Italy. PRO_COM codes are referred to 2025 administrative Italian municipalities borders. Informative spatial and temporal variables are: "PRO_COM": The ID code of the municipality "time": The day considered Then, the climate variables are the average in the municipality of the variables listed in table 2 (i.e. "mean_variable") along with the intra-municipal variability (i.e. "sd_variable"). Supplementary LAUs information Along with municipalities codes, geometries are available in the file named "". The two files are kept separate to save space. To retrieve geometries and merge with the LAUs datasets, an example with climate variables is available in the script "example_LAUs_merge_and_visualization". Tables Table 1: Air Quality variables present in the GRINS AQCLIM points dataset along with summary statistics. Unit of measure: micrograms per cubic meter. Air pollutant # stations Min. Mean Max. NA's (%) CO 268 0 0.5 58.2 34 NH3 1 0 5.71 17.26 57 NO2 710 0 21.71 362.21 27 NO 336 0 11.44 583.46 65 O3 402 0 56.45 756.33 29 PM10 648 0 24.08 2,575 30 PM2.5 351 0 15.51 907 35 SO2 274 0 3.11 868.42 36 Table 2: Climate variables in the GRINS_AQCLIM_points and in GRINS_CL_LAUs datasets Variable Name Description Unit CL_blh Daily mean of the height of the atmosphere boundary layer m CL_lai_hv Daily fixed value of high vegetation leaf area index m2/m2 CL_lai_lv Daily fixed value of low vegetation leaf area index m2/m2 CL_rh Daily mean of relative humidity % CL_ssr Daily maximum of surface solar radiation J/m2 CL_t2m Daily mean of temperature at 2 meters ∘C CL_tp Daily cumulative total precipitation m CL_winddir Daily mode of wind direction (1=N, 2=E, 3=S, 4=W) - CL_windspeed Daily mean of wind speed m/s



