遇见数据集

Data files and scripts for the manuscript entitled "Biogeographic regions within Indomalaya: an integrated approach using bird distributions"

收藏
Zenodo2026-01-22 更新2026-05-26 收录
官方服务:

资源简介:

This repository contains R scripts and data files associated with analyses presented in the manuscript entitled "Biogeographic regions within Indomalaya: an integrated approach using bird distributions" in the Journal of Biogeography (doi: 10.1111/jbi.70155). The presence-absence matrices can be used to replicate all three biogeographic region delimitation analyses presented in the manuscript using the code and instructions provided in the R scripts. Specific input files for each analysis are also provided. Users who wish to begin the process with raw unfiltered distribution maps should request the spatial data from BirdLife International (https://datazone.birdlife.org/). Description of the files (listed alphabetically): R scripts 01_IndomalayaBioregions_Dataset.R: This is the R script used for setting up and filtering the raw range distribution data into the final 1544-species dataset. This script also includes information about work done externally on ArcGIS and QGIS. 02_IndomalayaBioregions_NDM.R: This is the R script used for setting up the NDM input files (from the presence-absence matrix csv files), extracting information from output files from the NDM/VNDM software, and plotting the resulting AOEs as presented in the manuscript. This script also includes information about work done externally on NDM/VNDM. 03_IndomalayaBioregions_Ecostructure.R: This is the R script used for Ecostructure Grade of Membership model fitting, visualization of results as rasters, calculation of motif overlap, and construction of network graph tracking motif stability, as presented in the manuscript. 04_IndomalayaBioregions_hclust.R: This is the R script for the hierarchical clustering analysis, which includes procedures for selecting the best clustering algorithm and most optimal number(s) of clusters, as well as code for visualizing clustering results both on a dendrogram and on a map, as presented in the manuscript. Map spatial polygons Asia_Tropic_Map_2023.zip: Four files (.dbf, .prj, .shp, and .shx) associated with the shapefile for a rectangular map of tropical Asia that can be used as a basemap for plotting spatial data. indomalaya_map_2023.zip: Four files (.dbf, .prj, .shp, and .shx) associated with the shapefile for the Indomalayan Realm (as defined by the WWF) that is used as a reference for filtering range data and interpretations out all downstream output. Data and input files any_indomalaya.csv: This is the list of all 2827 species with any native and resident occurrence in Indomalaya. This list is the starting point for the area percentage filtering process in described in the script 01_IndomalayaBioregions_Dataset.R. Ecostructure_motif_names_all.zip: Four .csv files containing expanded motif names represented by three letter motif abbreviations shown in Figure 3 and Figure S3. The four filenames end in 5d, 2d, 1d, and hd, each for one of the four grid resolutions (5°, 2°, 1°, and 0.5° respectively). Each .csv also contains a column containing a series of “graphing indices” for positioning motif nodes when plotting the network graph. IM60_presab_matrices.zip: A .zip archive of four .csv files, each one a presence-absence matrix generated by the R package letsR for the final 1544-species dataset used in this study. “IM60” refers to “Indomalaya 60%.” Each filename ends in one of four strings: “5deg”, “2deg”, “1deg”, and “halfdeg”, for one of four analytical grid resolutions (5°, 2°, 1°, and 0.5° respectively). input_matrices.zip: A .zip archive of four .csv files, each one a presence-absence matrix of all 1544 study species, formatted to serve as the input file for Ecostructure model fitting, as well as a starting point for generating a beta-similarity distance matrices for Hierarchical Clustering. Each filename ends in one of four strings: “5deg”, “2deg”, “1deg”, and “halfdeg”, for one of four analytical grid resolutions (5°, 2°, 1°, and 0.5° respectively). Instructions for creating these matrices from the letsR output objects in an R environment can also be found in the R scripts on this data repository. NDM_input.zip: A .zip archive of four NDM input files. Each filename contains one of four stems: “5deg”, “2deg”, “1deg”, and “halfdeg”, for one of four analytical grid resolutions (5°, 2°, 1°, and 0.5° respectively). The files are in a .dat format and have been specifically formatted to be importable into the ndm/vndm software GUI. Analysis output files Ecostructure_XXdeg_modelfits.zip (XX = ‘5’, ‘2’, ‘1’, or ‘half’, for each of the four analytical grid resolutions): Four separate .zip archives, each containing all 19 Ecostructure model fit results from K = 2 to K = 20 in a .rds format. NDM_2percent_consensus.xlsx: This is a spreadsheet with four separate tabs (each corresponding to one of the four grid resolutions) that summarizes the NDM results contained in the .out files into one document. NDM_2percent_spList.zip: A .zip archive of four xlsx spreadsheets. Each filename contains one of four stems: “5deg”, “2deg”, “1deg”, and “halfdeg”, for one of four analytical grid resolutions (5°, 2°, 1°, and 0.5° respectively). Each spreadsheet lists the endemic avian species names for each 2% consensus areas found in by the NDM analysis. NDM_2percent_outfile.zip: A .zip archive of four NDM output files. Each filename contains one of four stems: “5deg”, “2deg”, “1deg”, and “halfdeg”, for one of four analytical grid resolutions (5°, 2°, 1°, and 0.5° respectively). The files are in a .out format and have been exported from the ndm/vndm software containing supplementary information about 2% consensus areas and the endemic species that support their delineation found at the given spatial resolution. Supplementary Figures Ecostructure_allrunsallmotifs.zip: Supplementary figures of all motifs identified from K=2 to K=20, with motif abbreviations presented in Figure 3 of the manuscript, across all four grid resolutions. Each individual map shows the distribution of grid cells that constitute each given motif. The transparency of the cells’ color is adjusted according to the assigned membership probability, or omega (ω) values, for the specific motif. The darker the shade (lower transparency), the higher the membership probability, and vice versa. hclust-UPGMA_maps.zip: Supplementary figures of all hierarchical clustering (UPGMA) results with dendrogram branch tips depicted as grid cells on a map of Indomalaya from K=5 to K=36 across all four grid resolutions, showing how UPGMA partitions Indomalaya into progressively finer biogeographic units. The color coding in all four figures is independent of that presented in Figure 4 of the manuscript. Also note that the areas not considered part of Indomalaya, e.g., the Korean Peninsula, Kyushu of Japan, Sulawesi, the Lesser Sundas, Western Pakistan, and Afghanistan, have been included in the analysis even after filtering, due to widespread Indomalayan species whose ranges exceeds the Indomalayan Realm boundaries. The species richness in these fringe areas had been artificially reduced prior to the analysis, so their appearances as well-delineated biogeographic areas at various points of the clustering analyses should be interpreted with caution.

提供机构:
Zenodo
创建时间:
2025-06-11
二维码
社区交流群
二维码
科研交流群
商业服务