遇见数据集

Data for "How should we identify dominant species? A multi-scale comparison of metrics in Amazonian forests"

收藏
Zenodo2026-09-28 更新2026-10-01 收录
官方服务:

资源简介:

Tree species dominance has been widely documented in tropical forests. Although various metrics have been used to identify dominant species, a methodological comparison across spatial scales is missing. Here, we compare four metrics for identifying dominant species (importance value index (IVI), F index, abundance, and h index) in tropical forests using a 140-plot network in western Amazonian terra firme forests. We compared among metrics (i) the number of dominant species identified, (ii) the overlap of dominant species, and (iii) the potential distribution and niche breadth of species identified as dominant by each metric. Iterative analyses across multiple data combinations and species niche modelling were employed to achieve these objectives. The F index and abundance displayed the highest similarities, with some discrepancies. The F index excludes as dominant those species that are highly abundant but restricted to a few plots. IVI exhibited notable differences, primarily attributed to the inclusion of basal area as a parameter for identification of dominant species. Finally, the h index identified the highest number of dominant species, exhibiting notable disparities in species composition, and biased towards dominant species with narrow niche breadth and lower potential distribution compared to the other metrics. The F index, followed by abundance, showed the greatest correspondence with the ecological characteristics of dominant species considered here, as it identifies oligarchic species with wide niche breadth and large potential distribution. The IVI proves particularly useful for studying dominance in relation to species abundance, biomass, and productivity. The h index was less aligned with the dominance criterion evaluated in this study because it dilutes differences in dominance among species and identifies such a large number of species as dominant that the resulting set exhibits narrower niche breadth and lower potential distribution than the other metrics. We set up a network of 140 0.1-ha (20 × 50 m) plots in western Amazonian terra firme forests, covering a latitudinal gradient of 1800 km, from Ecuador to Bolivia. Plots were located with a distance of at least 300 m apart, avoiding locations affected by human disturbances and natural gaps. We accounted for 1756 species. We compared four most used metrics for identifying dominant species (importance value index (IVI), F index, abundance, and h index). To test if metrics identify different number of dominant species and select different species as dominant across scales, we conducted the following iterative analysis: 1) We defined 28 range distances for plot selection within the network of 140 plots. These range distances were set at evenly spaced intervals, ranging from 55 km (the most local scale) to 1945 km (the most regional scale). 2) For each range distance, we started each iteration with one of the 140 plots at a time (i.e., origin plot). Then we selected all plots within the selected range distance, totaling 3920 iterations (140 origin plots x 28 range distances). 3) At each iteration, we calculated the number of species identified as dominant by each metric as well as the degree of concordance of dominant species between pairwise metrics. To check if significant differences existed in the number of dominant species identified by metric across scales, we used generalized linear mixed models (GLMMs) with a negative binomial error distribution. We used the number of dominant species by metric as the response variable and the metric (categorical) and range distance (continuous) as predictors. The most complex model included the interaction between both predictors. Origin plot (n =140) was used as random factor to account for potential correlation among sampled plots. Models were compared with the Akaike Information Criterion (AIC). Differences larger than 2 in the AIC of two models indicate that the model with a greater AIC was not supported and it was omitted. Model residuals were explored using a simulation-based approach to obtain readily interpretable scaled (quantile) residuals for the fitted GLMMs. As measure of goodness of fit, we used marginal R2 (R2m) to consider fixed effects and conditional R2 (R2c) to consider both random and fixed effects. Evaluation of metric performance To evaluate the performance of each metric in the selection of dominant species, we calculated the potential distribution and niche breadth of each dominant species by each metric at each range distances. To do so, we compiled the names of the dominant species detected across the 3920 iterations from prior analyses. Subsequently, we conducted ecological niche modelling using species occurrence data and environmental information from online databases, as follows: a) Species distribution data Morphospecies identified as dominant were removed because of the impossibility to obtain geolocated metrics. Thereby, we compiled species occurrences of 529 species from GBIF and BIEN when they corresponded to herbarium specimens with precise location information (no coordinate issues reported). To include records in the vast majority of the distribution area, as recommended for ecological analysis to avoid niche truncation, we obtained occurrences with coordinates for the bounding-box delimited by -82/-32 degrees of longitude and -38/12 degrees of latitude, encompassing the whole Amazon basin. We checked occurrences through automated evaluation to remove geographic and taxonomic errors. These automated processes included the spatial validation of recorded administrative information with administrative maps, with occurrences with a mismatch removed. Records identified as invalid, plain zeros coordinates, or least of two decimals precision were removed. Finally, we evaluated potential distribution area (by means of ecological niche models) and niche breadth for the 506 species with at least 15 unique presences. b) Environmental variables We considered both climatic and edaphic datasets as independent variables: the nineteen Chelsa 2.1 bioclimatic variables and the nine variables related to physical and chemical soil properties at 0-5 cm depth obtained from the SoilGrids 2.0; at a spatial resolution of 30 arc-seconds (≈ 1 km2 cell size). As background, we randomly selected 10,000 points over the entire study area. To avoid multicollinearity, we eliminated one of the variables in each pair with a Pearson’s correlation value > 0.8. We finally selected the variables isothermality (bio3), maximum temperature of warmest month (bio5), minimum temperature of coldest month (bio6), precipitation of driest quarter (bio17), precipitation of warmest quarter (bio18), precipitation of coldest quarter (bio19), bulk density (bdod), total nitrogen (nitrogen), pH in H2O solution (phh2o), sand content (sand), and soil organic carbon content (soc). c) Species potential distribution and ecological niche breadth We used ecological niche models (ENMs) to represent species potential distributions. We calibrated ENMs using the BIOMOD2 R package as ensembles of two statistical techniques: gradient boosting machine and random forest. We calibrated models with 70% of the data and evaluated with the remaining 30% using the area under the ROC curve (AUC). We replicated the procedure 10 times with random training and evaluation datasets. Eventually, we obtained 20 replicate models (10 replicates × 2 techniques). To remove spurious models, we generated the ensembles using the models with AUC > 0.8. The contribution of each model to the final ensemble model was proportional to their AUC value. We converted the ensemble model into a binary model (presence/absence) applying four different threshold criteria: a threshold allowing a maximum of 5% and 10% of omission error (i.e. omission error is the percentage of the real presence predicted as absences in the model); and the threshold maximizing AUC and the true skill statistic (TSS). The correlations were > 0.9 to all pairwise comparisons of the four thresholds. We therefore chose the AUC threshold for being the one that showed the highest correlation with the others. Given our work is focused on western Amazonia, we cut off the final potential area distribution as number of pixels of each dominant species to the unique extension of western Amazonia. We estimated ecological niche indexes per species. We performed a principal component analysis (PCA) of the environmental conditions recorded at 10,000 background points to identify the combinations of variables that best captured ecological variation. Then, we plotted the observations of each species and transformed them into densities in a gridded environmental space depicted by the first two PCA axes. Finally, we calculated niche breadth by calculating the Shannon index from environmental occupancy (the density of species occurrences in PCA space divided by the density of available environment and multiplied by its logarithm). d) Statistical analyses To detect differences regarding potential distribution and ecological niche breadth of dominant species selected by each metric across different scales, we conducted two groups of linear mixed models (LMMs) with different response variable: one with species mean potential distribution, and another with species mean ecological niche breadth. We used as predictors the metric (categorical) and the range distance (continuous). The most complex model included the interaction between both predictors. Origin plot (n = 140) was used as a random factor to account for potential correlation among sampled plots. Models were compared with AIC. As measure of goodness of fit we used R2m and R2c.

提供机构:
Zenodo
创建时间:
2026-09-28
二维码
社区交流群
二维码
科研交流群
商业服务