An open-source classifier for acoustic monitoring of hazel dormice: Annotated acoustic data
收藏资源简介:
Summary Acoustic data used to train an open-source classifier for hazel dormouse (Muscardinus avellanarius) vocalisations. The main set of field recordings were collected at Calke Abbey, Derbyshire, UK between 14 June and 1 July 2023. This data was annotated for hazel dormouse (Muscardinus avellanarius) vocalisations as part of a master's project. This dataset also includes supplementary recordings of UK bat species collected in southern England. These were included in the training dataset as negative samples encompassing common confusion species for hazel dormouse. Methodology The following is an extract from the methodology section of my master's project which explains how the dataset was collected and annotated. This also explains what each .zip file contains: Sample 1 (calke_abbey_sample1.zip), Sample 2 (calke_abbey_sample2.zip) and the additional bat field recordings (annotated_bat_data.zip). Please note the dataset shared here does not include annotations for bat species as these were provided to me in a different format - these were just used to confirm that the recordings contained only bats and no dormouse vocalisations. Data acquisition Acoustic data sources used for model training and validation comprised a primary dataset of field recordings taken in Derbyshire and supplementary field recordings obtained from existing datasets. The primary dataset was recorded between 14 June and 1 July 2023 at Calke Abbey, Derbyshire, when 38 hazel dormice were reintroduced to the site (PTES, 2025b). The reintroduction used a soft-release protocol (Resende et al., 2021). During the first 7 days of recordings (Phase 1), hazel dormice were held and provided with food in a temporary enclosure at the site with 2-3 dormice in each cage. During the next 7-day recording period (Phase 2), the enclosure was opened to allow the dormice to explore the site and they continued to receive supplementary feeding. Recordings were made using AudioMoth devices (Hill et al., 2019) in an IPX7 waterproof case positioned between 1cm and 40cm away from each cage. The AudioMoths recorded for 55 seconds each minute. The sample rate was set to 192kHz, gain was set to medium and no bandpass filter or frequency trigger was applied. There were 12 cages in total, although for seven cages acoustic recordings were only captured in either Phase 1 or Phase 2 due to equipment failure or devices being lost. Acoustic data was captured across both phases for the remaining five cages. Another set of field recordings collected during bat surveys in southern England provided examples of British bat calls which may be confused with dormice due to their similar ultrasonic range. These were annotated by a bat acoustics expert. Data annotation We selected data from three cages for annotation. Cages where recording had failed during either Phase 1 or 2 were excluded. AudioMoths were positioned at a range of distances from the cages (1cm, 18cm or 40cm) to account for the possible effect of distance on recording quality. The data were stratified according to cage and phase. Recordings were randomly sampled evenly from each strata, resulting in a sample of 360 55s recordings (Sample 1). We found that a very small proportion of files from Sample 1 contained positive dormouse annotations (7 of 360 55s recordings). Therefore, for Sample 2 a Goertzel filter (Goertzel, 1958) was applied to the data to identify recordings containing ultrasonic activity. We piloted the Goertzel filter on a small selection of positive recordings from Sample 1, trialling a range of different central frequencies across the dormouse’s frequency range (18-53kHz; Ancillotto et al., 2014). A central frequency of 42.5kHz performed sufficiently well, therefore we applied a Goertzel filter with a central frequency of 42.5kHz to the dataset. The primary dataset was filtered to a subset of recordings that were not included in Sample 1 and contained ultrasonic activity according to the Goertzel filter. To generate Sample 2, a further 360 55s recordings were selected from this subset using the same stratified sampling approach as in Sample 1. For Sample 1 and 2, hazel dormouse calls were inspected and annotated using Raven Lite (K. Lisa Yang Center for Conservation Bioacoustics at the Cornell Lab of Ornithology, 2023). Based on a pilot study in which 50 55s clips were annotated for hazel dormouse calls, two readily distinguishable call types were selected for annotation: rising calls (encompassing call types A, B and C described by Ancillotto et al. [2014]) and arch calls (part of the dormouse’s courtship song identified as call type D by Ancillotto et al. [2014]). Other hazel dormouse calls were annotated at species level without specifying the call type. Dataset structure File types Each .wav file is a 55-second AudioMoth recording. Each .wav file has a corresponding .txt annotation file containing boxed annotations produced in Raven, which lists all hazel dormouse vocalisations identified in the .wav file. Annotation codes Hazel dormouse annotations begin with 'hdor_', followed by the call type if identified. The annotations end in '_cam' if a hazel dormouse was observed on the camera trap located in the same cage as the audio recorder within approximately 10 minutes of the vocalisation or '_ncam' if not. Call type annotations refer to the hazel dormouse call types identified by Ancillotto et al. (2014), with codes listed below. Code Call type in Ancillotto et al. (2014) hdor_asc Call types A, B and C: Arising calls hdor_arch Call type D: Arched/downsweep calls hdor_wig Call type E: Flat wiggled calls hdor Dormouse, unspecified call type References Ancillotto, L., Sozio, G., Mortelliti, A. and Russo, D. (2014). ‘Ultrasonic communication in Gliridae (Rodentia): the hazel dormouse (Muscardinus avellanarius) as a case study’, Bioacoustics, 23(2), pp. 129–141. https://doi.org/10.1080/09524622.2013.838146 Goertzel, G. (1958) ‘An Algorithm for the Evaluation of Finite Trigonometric Series on JSTOR’, The American Mathematical Monthly, 65(1), pp. 34-35. https://doi.org/10.2307/2310304 Hill, A. P., Prince, P., Snaddon, J. L., Doncaster, C. P. and Rogers, A. (2019). ‘AudioMoth: A low-cost acoustic device for monitoring biodiversity and the environment’, HardwareX, 6, p. e00073. https://doi.org/10.1016/j.ohx.2019.e00073 K. Lisa Yang Center for Conservation Bioacoustics. (2023). ‘Raven Lite: interactive sound analysis software (2.0.4)’, Cornell Lab of Ornithology. Available at: https://ravensoundsoftware.com/ People’s Trust for Endangered Species. (2025b). Press release: First ever reintroduction of rare hazel dormice into the National Forest. Available at: https://ptes.org/press-release-first-ever-reintroduction-of-rare-hazel-dormice-into-the-national-forest/ Resende, P. S., Viana-Junior, A. B., Young, R. B. and Azevedo, C. S. (2021). ‘What is better for animal conservation translocation programmes: Soft- or hard-release? A phylogenetic meta-analytical approach’, Journal of Applied Ecology, 58(6), pp. 1122–1132. https://doi.org/10.1111/1365-2664.13873



