Mapping Genuine Wild Olive Trees in the Mediterranean Basin: Computational and Reproducibility Archive
收藏资源简介:
1. Overview This repository contains the computational materials associated with the Master's thesis: A Geospatial and Machine Learning Framework for Mapping Genuine Wild Olive Trees: Distinguishing Genuine Wild, Admixed, and Cultivated Occurrences in the Mediterranean Basin The archive provides the scripts, input datasets, intermediate computational products, model outputs, figures, tables, GIS maps, and reproducibility documentation required to document the analytical workflow used in the thesis. The accompanying reproducibility document is entitled: Reproducibility Workflow and Computational Analysis It documents the Google Earth Engine (GEE) predictor-extraction procedures and the R-based analyses used for Model 0, Model 1, the Sensitivity Analysis, Model 2, and application of the final model to the curated GBIF occurrence dataset. The workflow distinguishes three olive-status classes: Genuine Wild Admixed Cultivated The archive is organized to preserve the chronological and computational logic of the analysis rather than to provide a single monolithic executable workflow. 2. Reproducibility Scope The computational workflow combines: GBIF occurrence-data curation; literature-derived biological reference labels; Google Earth Engine environmental predictor extraction; anthropogenic and landscape-context predictor extraction; spatial data processing; spatial cross-validation; Random Forest classification; predictor selection; model tuning; model-performance evaluation; sensitivity analysis; confirmatory model comparison; final Model 2 fitting; interpretation and visualization; application of the final model to the curated GBIF dataset. The workflow was designed to maintain a strict distinction between the labelled reference dataset used for model development and the unlabelled curated GBIF dataset used as the target prediction dataset. The literature-derived reference observations provide the class labels used for model development. The curated GBIF dataset is treated as the prediction target and is not used to create the biological class labels.



