遇见数据集

Global particulate pollution research reflects economic capacity more than pollution burden

收藏
Zenodo2026-07-01 更新2026-08-01 收录
官方服务:

资源简介:

This repository contains the data, code, and figure outputs supporting the study "Global particulate pollution research reflects economic capacity more than pollution burden" (Yang et al.). The work analyzes more than 55,000 PM2.5-related scientific abstracts published between 1980 and 2025 (Web of Science) to map the global distribution of air-pollution research and compare it against aerosol burden and socioeconomic need. Methods reproduced here. A large language model (OpenAI GPT-4o-mini, Chat Completions API) is used to geoparse study locations from abstracts onto a 1° × 1° global grid (geoparsing F1 ≈ 0.75, bootstrapped). The resulting publication-density field is combined with MERRA-2 aerosol optical thickness (AOT, 1980–2025), WHO health burden indicators (DALYs, PM2.5-attributable death rate), and World Bank socioeconomic data (GDP, Gini index, population), all regridded to a common 1° grid. An interpretable XGBoost regression with SHAP analysis (10-fold cross-validation) attributes research disparities to their underlying drivers, identifying GDP and population — rather than aerosol or health burden — as the dominant predictors of research intensity (the "knowledge–exposure gap"). Contents. Code: PM2.5_LLM.py (LLM geoparsing pipeline with the fixed 1°-grid prompt template), Machine Learning.py (XGBoost 10-fold CV + SHAP), and Figures.ipynb (figure generation), plus notebook equivalents. Literature corpus: cleaned PM2.5 bibliographic records and the structured LLM geoparsing outputs (region names, bounding boxes, central points) with parse audit logs. Gridded products: 1° publication-density fields (NetCDF/CSV), annual and mean MERRA-2 AOT fields, and the machine-learning input table (xgb_inputs_1deg_final.csv). Results: cross-validation metrics and predictions, feature-importance and SHAP figures, knowledge–exposure gap index, spatial-resolution sensitivity analyses (1°/5°/10°), regional time series, and publication–AOT overlay maps. Data sources. MERRA-2 reanalysis (NASA GMAO); WHO Global Health Estimates; World Bank Open Data (GDP, Gini, population). LLM geoparsing used the gpt-4o-mini model with a fixed prompt template and consistent parsing rules; API keys are not included in this repository.

提供机构:
Zenodo
创建时间:
2025-10-26
二维码
社区交流群
二维码
科研交流群
商业服务