The MESOSCAN Dataset: A Collection of Mesofauna Images from UK Forest and Peatland Soils
收藏资源简介:
A UK forest and peatland soil mesofauna image dataset collected by Forest Research from 2024 to 2025 onwards. Mesofauna are arthropods that live in the soil ranging from approximately 2mm down to 0.2mm in size. Their biodiversity is a key indicator to the health and sustainability of the soil. This dataset is a systematic attempt at collecting and curating valuable images that can be used to help recognise specimens, and to provide training data to machine learning systems that can substantially automate identification and make monitoring more scalable. An example, baseline, machine learning model is provided to validate the data and to illustrate how it can be used - see the ./notebooks folder for this. Highlights: Currently 2600+ training images across 35 taxa each 1600x1600 pixels. Samples were currated from 21 sites across the UK with specimens taxonmically identified prior to imaging. Emphasis has been on creating a varied and extensive range of views for a number of specimens. A selection of realistic and independent test image are provided from cluttered assemblages. Structure The dataset is principally focused on mites (Acari) and springtails (Collembola) which together make up of the majority of mesofauna biomass of soil samples. It organises the data within a (simplified) taxonomic structure reflected by the following directory levels contained in the ./train folder: Class Group Family Genus Species Acari Mesostigmata Trachytidae Trachytes aegrota Acari Mesostigmata Uropodidae Uropoda minima Acari Mesostigmata Rhodacaridae Rhodacarellus silesiacus Acari Mesostigmata Parasitidae Paragamasus Acari Mesostigmata Veigaiidae Veigaia nemorensis Acari Astigmata (_nymphs) Acari Oribatida Galumnidae Acrogalumna longipluma Acari Oribatida Scheloribatidae Liebstadia similis Acari Oribatida Phthiracaridae Steganacarus michaeli Acari Oribatida Phthiracaridae Phthiracarus silesiacus Acari Oribatida Brachychthoniidae Sellnickochthonius zelawaiensis Acari Oribatida Nanhermanniidae Nanhermannia coronata Acari Oribatida Nanhermanniidae Nanhermannia coronata (_nymphs) Acari Oribatida Malaconothridae Malaconothrus monodactylus Acari Oribatida Euzetidae Euzetes globulus Acari Oribatida Hypochthoniidae Hypochthonius rufulus Acari Oribatida Euphthiracaridae Acrotritia duplicata Acari Oribatida Oppiidae Oppiella nova Acari Oribatida Quadroppiidae Quadroppia Acari Oribatida Camisiidae Platynothrus peltifer Acari Oribatida Camisiidae Platynothrus peltifer (_nymphs) Acari Oribatida Achipteriidae Parachipteria punctata Acari Oribatida Peloppiidae Ceratoppia bipilis Acari Oribatida Ceratozetidae Ceratozetes gracilis Acari Oribatida Nothridae Nothrus pratensis Acari Oribatida Nothridae Nothrus pratensis (_nymphs) Acari Endeostigmata Nanorchestidae Nanorchestes Acari Prostigmata Eupodidae Collembola Entomobryomorpha Isotomidae Folsomia penicula Collembola Entomobryomorpha Isotomidae Isotomiella minor Collembola Entomobryomorpha Isotomidae Parisotoma notabilis Collembola Entomobryomorpha Entomobryidae Lepidocyrtus lignorum Collembola Poduromorpha Neanuridae Friesea truncata Collembola Neelipleona Neelidae Megalothorax minimus Collembola Symphypleona Sminthurididae Sminthurides malmgreni At any level within this folder structure images can be stored - according to the taxonomic level of identification that has been resolved. For example, where exact species have been identified the images taken will be found in the very specific folder for that species, whereas more generic identification will be found further up the hierarchy. When specimens cannot be resolved - or else form morphotypes - then a subfolder (begining with an _ underscore) will be created at that point in the taxonomy (e.g. _nymphs). Contents Alongside this README.md file (with example image mesoscan.png) plus the LICENCE file, this dataset contains: MESOSCAN.csv - Top level Excel/OpenOffice spreadsheet /train - Top level folder containing training images /scripts - Some utility Python scripts /calibration - Images taken prior to data processing /notebooks - Jupyter python code demonstrating and analysing data The MESOSCAN.csv file describes the source and nature of all the samples used and how processed, as follows: PROCESS_ID - unique ID (date time + sequence on day) of when images captured SAMPLE_PROJECT - record of project where sample was collected from SAMPLE_SITE - unique ID of site from which sample taken SAMPLE_HABITAT - description of site habitat SAMPLE_SOIL_TYPE - description of site soil type SAMPLE_DATE - date on which original sample was take from site SAMPLE_LAT - latitude (WGS84) of site SAMPLE_LON - longitude (WGS84) of site TAXON_ID - unique coding of sample specimens TAXA_CLASS - simplified top level taxonomic rank TAXA_GROUP - simplified next level taxnommic rank TAXA_FAMILY - family of specimens TAXA_GENUS - genus of specimens TAXA_SPECIES - species of specimens FORM_ATTRIBUTE - additional type of specimens NUM_INDIVIDUALS - bumber of individual specimens used from sample NUM_IMAGES - number of images captured on this processing run CURATOR - initials of data capture operator and curator The following subsections describe each of the subfolder contents. Train Full original images were captured at a resolution of 6960 x 4640 located when moving across a sample (from origin bottom left) moving in negative X across (left to right), negative Y up (bottom to top), and negative Z (down towards sample). Each specimen image was found and cropped to a resolution of 1600 x 1600 pixels and is is uniquely named with a format that describes: The date (also process ID) and time of capture run - e.g. 241106_110952 (on the 6th of November, 2024 at 11:09:52) The X location in the sample of the camera position - e.g. _X-7_300 (-7.3mm from origin in X) The Y location in the sample of the camera position - e.g. _Y-27_400 (-27.4mm from origin in Y) The Z depth location of the camera position - e.g. _Z-46_920 (-46.92 from origin in Z) The pixel X and Y origin position in the original image to crop from - e.g. _860_1134 Particular attention has been given to capturing a number of high resolution images with multiple views of multiple specimens including at multiple depths. Calibration The full resolution images in this folder were taken - as indicated by the date/process_id - of a microscope calibration slide indicating 1 mm central ruled scale. This approximates 100μm ≈ 125 px. All images processed between the dates of calibration were taken with that. Scripts The ./scripts folder has a selection of useful Python programs as follows: check_all_files.py - checks and summarises dataset files. prepare_for_tf.py - creates a "flat" training folder with symobolic links to data, suitable for loading into TensorFlow. Notebooks The following Jupyter notebooks are used to build a baseline model, test the model on more complex images, and demonstrate/plot further analysis or visualisation: mesoscan-build - main notebook illustrating construction of model mesoscan-sites - plots map of site locations NOTE as well there is a python/pip requirements.txt file that specifies dependencies for the jupyter kernel. Funding and Licencing Data collection was funded under the UK Governments Natural Capital Environment Assessment (NCEA) programme and the Department for Energy, Security and Net Zero (DESNZ) for sites in Scotland. Additionally, some work was supported by funding made available from the Scottish, Welsh, and UK government via the Science and Innovation Strategy for Forestry in Great Britain (2021-2025), the Nature for Climate Fund (NCF) Programme, and Forest Research and Development (FRD) Programme. The data is Crown Copyright and released under the Open Government Licence (v3).



