遇见数据集

Multi-resolution sampling-effort data for global biodiversity modelling

收藏
Zenodo2026-05-10 更新2026-05-26 收录
官方服务:

资源简介:

This Zenodo repository provides an archival distribution of the datasets generated in the study: El-Gabbas, A. (2026). A global, taxon-stratified, high-resolution sampling-effort dataset from GBIF for bias-aware ecological modelling. Diversity and Distributions. https://doi.org/10.1111/ddi.70205 The dataset was derived from billions of biodiversity occurrence records available through the Global Biodiversity Information Facility (GBIF) using a fully reproducible R-based workflow. The resulting products consist of global raster layers representing biodiversity sampling effort using two complementary metrics: observation count (number of occurrence records per grid cell) and species richness (number of unique species per grid cell). Data are provided across nine major taxonomic groups and their descendants, multiple spatial resolutions (~1, 5, 10, and 20 km), and both annual (1980–2025) and cumulative temporal products. This Zenodo deposition is intended primarily as a stable archival and reference distribution of the final processed outputs associated with the publication. To facilitate direct access and offline use, data are organised into compressed archives by major taxonomic group, with each archive containing the complete set of raster products for that group in a ready-to-download format. However, downloading entire archives is generally not the recommended workflow for most applications. The complete dataset is extremely large, and full taxonomic downloads may require transferring many gigabytes of data, most of which are often unnecessary for a specific analysis. Instead, users are strongly encouraged to access the dataset programmatically through the dedicated functions implemented in the ecokit R package together with the project GitHub repository: https://github.com/elgabbas/global_sampling_efforts. The GitHub repository provides complete documentation, reproducible workflows, usage examples, and guidance for efficient data access. The associated ecokit functions allow users to selectively retrieve only the specific files required for a given analysis, including particular taxonomic groups, spatial resolutions, years, or cumulative products, thereby substantially reducing download size, storage requirements, and local data-management overhead. The primary cloud-hosted distribution of the dataset is available through the Open Science Framework (OSF): https://osf.io/hz4sy. The OSF repository serves as the backend storage used by ecokit for selective and on-demand downloads. Rather than downloading complete archives, users can therefore retrieve only the exact raster products needed for their analyses, which is the preferred and recommended approach for most ecological and modelling workflows. The dataset is designed to support bias-aware ecological analyses, particularly species distribution modelling (SDMs), macroecological research, biodiversity monitoring, conservation planning, and the identification of spatial and temporal knowledge gaps in biodiversity data. If you use these data, please cite the associated paper. ------------------------------------------------------------------------------------------------------------------------------------------ Contents This dataset includes multi-resolution GeoTIFF raster layers representing two complementary metrics: Observation density (n_obs) – number of cleaned occurrence records per grid cell Survey species richness (n_sp) – number of unique GBIF speciesKey values per grid cell (excluding non-species-level identifications) Both metrics are provided as: Annual rasters for each year from 1980 to 2025, and Cumulative (“total”) rasters combining all records since 1980. Each product is available at four spatial grains: ~1 km (30 arc-seconds) ~5 km (2.5 arc-minutes) ~10 km (5 arc-minutes) ~20 km (10 arc-minutes) Taxonomic scope Nine major taxonomic groups are covered: Amphibia, Arachnida, Aves, Fungi, Insecta, Mammalia, Mollusca, Reptilia, and Tracheophyta (vascular plants). For each group, additional rasters are produced for descendant-level taxa at the next standard rank (e.g. orders within classes, classes within phyla), following the hierarchy described in Table 1 of the manuscript. Additionally, the total number of species and records per grid cell is also provided: "total.zip". Directory and file structure For each group, raster files are organised by metric and spatial grain using the following directory names: res_1_n_obs/ res_1_n_sp/ res_5_n_obs/ res_5_n_sp/ res_10_n_obs/ res_10_n_sp/ res_20_n_obs/ res_20_n_sp/ Each directory contains annual and cumulative GeoTIFFs named according to the convention: <metric>_<group>_<time>_res<resolution>.tif Examples:n_obs_Insecta_2020_res_5.tif / n_sp_Passeriformes_total_res_1.tif All rasters use the WGS 84 coordinate reference system (EPSG:4326) and ZSTD compression. Citation Users should cite the accompanying data paper as follows: El-Gabbas, A. (2026). A global, taxon-stratified, high-resolution sampling-effort dataset from GBIF for bias-aware ecological modelling. Diversity and Distributions. https://doi.org/10.1111/ddi.70205

提供机构:
Zenodo
创建时间:
2026-05-06
二维码
社区交流群
二维码
科研交流群
商业服务