phyloSDM_MOL data archive: inputs and outputs for phylogenetic species distribution models
收藏资源简介:
This archive contains the input, intermediate, and output data for the phyloSDM_MOL pipeline, which fits phylogenetically-informed species distribution models — a Log-Gaussian Cox Process (LGCP) with a phylogenetic Gaussian Process prior on species' environmental-response coefficients — to amphibian occurrence, expert range, and phylogenetic data. Closely related species are modeled as sharing similar environmental responses, allowing the model to borrow statistical strength across the phylogeny, particularly for data-poor species. The dataset covers 60 amphibian species across 14 genera (Ranidae and Hyperoliidae, including Chalcorana, Kassina, Hyperolius, Amnirana, Hydrophylax, Phlyctimantis, Papurana, Pulchrana, Sylvirana, Humerana, Hylarana, Abavorana, Acanthixalus, and Semnodactylus), organized into phylogenetically coherent clusters for model fitting. This archive is intended to be used together with the analysis code, maintained separately on GitHub: https://github.com/shubhi124081/phyloSDM_MOL. See that repository's README for full pipeline documentation, environment setup, and run instructions. Contents: raw_data/ — the harmonized and phylogeny-pruned working species list, per-species occurrence records (harmonized against Map of Life taxonomy), species-to-cluster assignments, and per-cluster intermediate files (train/test splits, packaged model input data) produced by the early pipeline stages. expert_ranges/ — expert-drawn range maps (GeoPackage, one per species) plus a combined shapefile, used to construct training/background data and prediction-time range priors ("soft clips"). analysis/ — pipeline outputs: per-species soft-clip rasters, phylogenetic conditional predictions, model evaluation results, continuous relative-probability-of-occurrence rasters, and thresholded binary range predictions. jobs/, log/ — example SLURM/dSQ job scripts, generated job-array task lists, and job manifests/logs from the HPC runs this pipeline was originally executed with (Yale McCleary cluster). data/, res/ — empty directories reserved for user-generated intermediate and model-fit files when re-running the pipeline; included for directory-structure completeness. Usage: Clone the GitHub repository above, then extract this archive into the repository root so that raw_data/, expert_ranges/, analysis/, jobs/, log/, data/, and res/ sit alongside the scripts/ directory from the code repository. Note: the pipeline additionally requires a set of global environmental raster layers (CHELSA bioclimatic variables, cloud cover, EVI, topographic ruggedness, elevation) that are not included in this archive due to size. See the GitHub README for details.



