Modelled breeding distributions of British birds at 1-km resolution
收藏资源简介:
This dataset provides modelled breeding-season distributions for 142 bird species across Great Britain at a 1 km × 1 km resolution. It is structured as a grid-cell-by-species table, with one row for each occupied 1 km × 1 km British National Grid cell and one binary occurrence column for each species. The first three columns identify the spatial unit: os_gridref gives the 1 km × 1 km Ordnance Survey grid reference, while easting and northing provide the associated British National Grid coordinates. Species columns are named using scientific binomials with spaces replaced by underscores. Values in the species columns are binary. A value of 1 indicates that the grid cell was included within the estimated breeding distribution of the species, while 0 indicates that it was not included. Grid cells for which no species was predicted to occur were excluded from the compiled dataset. The dataset therefore represents estimated distributions rather than direct observations and is intended for regional and national-scale biodiversity assessment, spatial prioritisation and related analyses. Individual cell values should not be interpreted as confirmation that a species is currently present at a particular location. Data sources and preparation Species distribution models were developed using eBird (Sullivan et al. 2009) complete-checklist data collected across Great Britain during the breeding season, defined here as March to August, between 2000 and 2022. Complete checklists report all bird species detected during a survey and were therefore interpreted as detection–non-detection data. A species was classified as detected when it was recorded on a checklist and as not detected when it was absent from an otherwise complete checklist. Surveys lasting more than 120 minutes or covering more than 5 km were excluded. To reduce the influence of repeated sampling in well-surveyed locations, records were thinned within combinations of 1 km × 1 km grid cell, observation coordinates, year, week and occurrence status. Environmental predictors were assembled at a 1 km × 1 km resolution. Land-cover predictors were derived from the UKCEH Land Cover Map 2021 (Marston et al. 2022) and represented the area of each land-cover class within each grid cell. Species-specific sets of land-cover variables were selected using independently defined breeding-habitat associations, together with a common set of major land-cover classes retained across species. Additional landscape variables included the lengths of hedgerows and rivers within each grid cell. Macroclimatic conditions were represented using HadUK-Grid climate data (Met Office 2018). These variables described seasonal and annual climatological conditions for the 1991–2020 reference period and included annual rainfall, mean annual temperature, seasonal growing conditions, and winter minimum-temperature or frost-related conditions. Together, the land-cover and climate predictors characterised broad spatial variation in habitat availability and environmental suitability across Great Britain. Variables describing survey timing and effort were included to account for differences in the probability of detecting a species among visits. These included survey week, year and hour, checklist duration, distance travelled and survey protocol. Easting and northing of the survey location were also included to capture broad-scale spatial structure not fully represented by the measured environmental predictors. Species distribution modelling Following the general approach of Johnston et al. (2021), separate calibrated random-forest models were fitted for each species. The random forest related observed detection–non-detection records to environmental conditions, survey timing and effort, and spatial location. Models used probability forests with an extremely randomised-tree splitting rule and class weights that increased the contribution of the less frequent detection records. Models were evaluated using spatial block cross-validation, in which records from each spatial block were withheld in turn. Blocks containing fewer than five detections were not treated as independent validation folds, and models were only fitted where at least two suitable blocks and 25 species detections were available. This spatial partitioning provided a more realistic assessment of the ability of each model to predict into locations that were geographically separated from the observations used for model fitting. The median (lower and upper 95% quantiles) across all species were 0.86 (0.66, 0.96). Individual species AUC summaries are included in auc_cv_for_breeding_birds_1km_great_britain_version_1.csv. Predicted probabilities from each cross-validation model were calibrated against the observed detection–non-detection data using a monotonic, shape-constrained generalised additive model. This calibration step improved the relationship between raw random-forest predictions and observed occurrence frequencies while ensuring that calibrated occurrence probability increased monotonically with the original model prediction. Models were projected across the 1 km × 1 km environmental grid under standardised survey conditions. Holding the timing, effort and protocol variables at representative values reduced the extent to which mapped predictions reflected spatial variation in observer behaviour rather than the underlying distribution of each species. Conversion to binary distributions A separate spatial projection was generated from each cross-validation fold. Continuous calibrated occurrence probabilities were converted to binary predictions using a species- and fold-specific threshold corresponding to 80% sensitivity. At this threshold, the model correctly classified approximately 80% of observed detections. This criterion was selected to provide a reasonable balance between omission error, in which genuinely occupied locations are excluded, and commission error, in which occurrence is predicted across an unrealistically broad area. The binary predictions from the cross-validation folds were then combined into a consensus distribution. A grid cell was classified as occupied where more than five fold-specific models predicted occurrence. This consensus step reduced the influence of individual spatial partitions and retained areas that were consistently predicted as suitable across multiple model fits. The resulting maps therefore represent consensus, binary estimates of recent breeding-season distribution at a 1 km × 1 km resolution. They integrate fine-resolution relationships with land cover and climate, adjustment for variation in survey timing and effort, spatially blocked model evaluation, and probability calibration. The R code used to develop these modelled species distributions is available at https://github.com/davidjbaker79/gbSDM.git. References Johnston, A., Hochachka, W.M., Strimas‐Mackey, M.E., Ruiz Gutierrez, V., Robinson, O.J., Miller, E.T., Auer, T., Kelling, S.T. and Fink, D. (2021). Analytical guidelines to increase the value of community science data: An example using eBird data to estimate species distributions. Diversity and distributions, 27(7), pp.1265-1277. Marston, C., Rowland, C.S., O’Neil, A.W., Morton, R.D. (2022). Land Cover Map 2021 (10m classified pixels, GB). NERC EDS Environmental Information Data Centre. https://doi.org/10.5285/a22baa7c-5809-4a02-87e0-3cf87d4e223a Met Office, Hollis, D., McCarthy, M., Kendon, M., Legg, T., Simpson, I. (2018). HadUK-Grid gridded and regional average climate observations for the UK. Centre for Environmental Data Analysis, 2023. http://catalogue.ceda.ac.uk/uuid/4dc8450d889a491ebb20e724debe2dfb/ Sullivan, B.L., Wood, C.L., Iliff, M.J., Bonney, R.E., Fink, D. and Kelling, S. (2009). eBird: A citizen-based bird observation network in the biological sciences. Biological conservation, 142(10), pp.2282-2292.



