遇见数据集

Data and code for: Physical environmental indicators of protected area representativeness: a machine-learning approach across 35 Japanese national parks

收藏
Zenodo2026-08-06 更新2026-08-13 收录
官方服务:

资源简介:

This dataset and code accompany the manuscript: Ise, H. & Kamada, M. Physical environmental indicators of protected area representativeness: a machine-learning approach across 35 Japanese national parks. Ecological Indicators (under review). The deposit provides a machine-learning pipeline for nationwide mapping of physical environmental representativeness in protected area networks, demonstrated across all 35 Japanese national parks. The indicator, PhysRep (Physical Representativeness score), is a continuous measure of how closely physical environmental conditions at any location resemble those within a park's strictly protected zones. Park-specific XGBoost classifiers are trained on climatic, topographic, geological and micro-landform variables within an H3 hexagonal grid at approximately 0.1 km² resolution, evaluated using spatially grouped nested cross-validation, interpreted with SHAP values, and extrapolated nationwide to produce continuous similarity surfaces. What is new in this version Corrected source elevation data. A coverage gap in the source DEM affecting the Yaeyama Islands region (Iriomote-Ishigaki and part of Ogasawara) had caused elevation and slope values for those cells to fall back to the nationwide mean during imputation. This has been corrected, and all 35 park-specific models have been retrained on the corrected dataset. All model outputs in this version supersede those in previous versions. Negative control validation. A label-shuffled negative control model is now included for every park, trained with an otherwise identical pipeline. Every park-specific model outperformed its own control in ROC-AUC. Control metadata for all 35 parks is included in the code archive under data/results/{slug}_shuffle_g8_gap1/. Contents japan-park-env-similarity_code_v3.zip — full codebase (preprocessing, modeling, analysis and figure notebooks, DuckDB SQL scripts, configuration), together with the per-park model metadata JSON files for both the production and negative control runs japan-park-env-similarity_data_v3-1.zip … japan-park-env-similarity_data_v3-8.zip — processed input datasets, trained models, and nationwide similarity outputs Due to file size limitations the data are distributed across eight archive files. Download all parts and extract them into the repository root directory; this restores the data/ directory structure expected by the notebooks. Data files File Description h3_jpn_res9_source_imputed.parquet Merged and imputed environmental dataset for all H3 cells across Japan h3_jpn_res9_processed.parquet Preprocessed dataset with English park name slugs and one-hot encoded categorical variables; direct input to the modeling notebook {slug}_similarity_scores_cv.csv H3 cell IDs with PhysRep scores for all of Japan {slug}_xgb_model_cv.joblib Trained XGBoost model {slug}_nationwide_similarity_cv.parquet Full nationwide dataset with PhysRep scores and binary labels Reproducibility Model outputs may differ slightly from those reported in the manuscript if the pipeline is re-run, due to stochastic elements in the negative sampling procedure, and this applies to the negative controls as well, where the label shuffle is an additional source of run-to-run variation. The values reported in the manuscript were obtained from the output files deposited here. See README.md in the code archive for the full workflow. Code is released under the MIT License; data outputs under CC BY 4.0.

提供机构:
Zenodo
创建时间:
2026-08-06
二维码
社区交流群
二维码
科研交流群
商业服务