遇见数据集

SAGE: A sampling-aware global evaluation benchmark for species distribution modeling - trained models and predictions

收藏
Zenodo2026-08-25 更新2026-10-01 收录
官方服务:

资源简介:

Trained models, per-plot test predictions, and per-species evaluation metrics for the main-table models of the SAGE (Sampling-Aware Global Evaluation) benchmark. SAGE evaluates species distribution models (SDMs) across 5,771 non-anonymized plant species by relating performance to the data conditions under which each species is observed. This record accompanies, and is a supplement to, the SAGE data record and the code repository. Use it to reproduce the paper's main results without retraining, or to load a trained deep model directly. The record covers the main-table models only; ablation runs are reproducible from the code, configs, and seeds in the repository. Artifacts are bundled by type into four tarballs, each expanding to per-run folders named <model>_s<seed>: - predictions.tar — per-plot test predictions (probabilities, 42,268 plots × 5,771 species) for all 50 main-table runs: 25 deep-model and 25 single-species runs, 5 seeds each.- checkpoints.tar — one trained PyTorch Lightning checkpoint per deep-model run (the best-validation-AUROC checkpoint, i.e. the one used to produce the reported predictions), 25 runs.- configs.tar — the resolved run configuration for each deep-model run, 25 runs. Required alongside a checkpoint: the checkpoint stores weights, and config.yaml is what rebuilds the network around them.- metrics.tar — per-species test AUROC and AUPRG for all 50 runs (25 deep and 25 traditional single-species models: MaxEnt, Random Forest, BRT, GLM, GAM). Two further files sit at the top level: - eval_test_targets.h5 — the shared ground truth (targets, plot locations, surveyed area) that every prediction file aligns to by row index.- sage-sdm-benchmark_code.zip — an archived snapshot of the code repository at the commit these runs were produced with. The single-species baselines are released as per-plot predictions and metrics; their fitted model objects are not distributed, being reproducible from the data, code, and seed. The 5,771 species follow the canonical order in the data record. Full layout, the configuration and run-ID mapping, integrity checksums, and a metric-recomputation example are documented in the included README.md.

提供机构:
Zenodo
创建时间:
2026-08-25
二维码
社区交流群
二维码
科研交流群
商业服务