BENTH-MARC: Georeferenced zoobenthic biodiversity data from the marine Arctic
收藏资源简介:
Abstract As the Arctic Ocean undergoes rapid environmental change and is subjected to increasing anthropogenic pressures, there is a growing need for comprehensive baseline data on its biodiversity. Although sampling efforts have been considerable over the past decades, the existing biodiversity records are often fragmented and scattered across multiple databases. We present BENTH-MARC, an open pan-Arctic database with BENTHos in the Marine ARCtic. BENTH-MARC contains 2,812,548 georeferenced zoobenthic taxon records of 8,331 species and 3,285 taxa at other taxonomic levels from 684 datasets and 210,867 stations. Data span across the entire pan-Arctic realm and integrate regional and global biodiversity databases as well as previously unpublished datasets. The records in BENTH-MARC come from two and a half centuries of sampling, from 1773-2026, and are based on morphological identification. By providing harmonized and accessible data, BENTH-MARC can support the assessment of biological consequences of ongoing environmental change, aid in the identification of biodiversity hotspots and evaluation of conservation targets. The database can also be used to guide future sampling efforts. Data sources BENTH-MARC includes publicly available, and previously unpublished, georeferenced and morphologically identified metazoan datasets of meio-, macro- and megabenthos in the Arctic LMEs. Only invertebrates were included, more specifically from the following phyla: Annelida, Mollusca, Arthropoda, Echinodermata, Cnidaria, Porifera, Sipuncula, Bryozoa, Chordata, Brachiopoda, Nematoda, Nemertea, Tardigrada, Platyhelminthes, Hemichordata, Kinorhyncha, Priapulida, Loricifera, Phoronida and Entoprocta. Data were extracted from GBIF and OBIS on 2026-05-05. The records were downloaded using the R packages rGBIF and robis, which offer programmatic access to the application programming interface (API) of GBIF and OBIS. To access associated metadata and information, complementary downloads targeted the related Extended Measurement or Fact (eMoF) extensions. The GBIF downloads are available at at https://doi.org/10.15468/dd.a8vs9g26 and https://doi.org/10.15468/dd.qvy3m627. Data from CRITTERBASE were downloaded through Alfred Wegener’s Institute’s dedicated web portal on 2026-02-06. Data from OGSL were obtained from the laboratory of P. Archambault (personal communication, 11 March 2026). References to all included datasets are available in the datasets.csv file. Data availability The data is primarily available as two separate files; one Parquet file with all occurrence records in a long format, and one CSV file with dataset-level metadata. The column datasetKey links these two files. Parquet is a column-oriented data storage format that facilitates efficient data processing and retrieval of big data. Unlike row-based formats like CSV, Parquet organizes data by columns, which allows for better compression and faster analytical queries of big datasets. Although there are no strict upper limits to how much data you can store in a CSV file, analyzing large CSV files requires significantly more processing power than Parquet files. To accommodate users who cannot access the data in Parquet format, the occurrence records are also available as one CSV file per Arctic Large Marine Ecosystem. All files are available on Zenodo at https://doi.org/10.5281/zenodo.20782806. Usage license is under Creative Commons Attribution 4.0 International (CC-BY-4.0). Code availability The R, Python and LaTeX code that was used to create the database BENTH-MARC and to generate this manuscript can be found in the GitHub repository benth-marc at https://github.com/hanna-modin/benth-marc. The workflow behind the database includes 19 scripts, of which five are in R and 14 in Python. Note that certain scripts are computationally intensive and require substantial CPU and memory resources. In the main folder of the GitHub repository and on Zenodo are also Excel files that give an overview of the workflow, glossary, mapping of columns from source databases to BENTH-MARC, and the main taxonomic filter. In workflow.xlsx, we have described the execution order, purpose, location, potential input and output files, and potential requirement of manual work associated with each script. The full process can be repeated without any new manual work if the same input occurrence data is used, but if new harvests from OBIS and GBIF are performed, there will be a need to manually update the filters and eMoF selection based on the new datasets. In glossary.xlsx, there is an overview of all 67 columns in BENTH-MARC's main occurrence file and the 12 columns in the dataset-level file. We have specified the data types and data coverage of each column, added explanations, and marked which columns are valid Darwin Core terms. The file column_mapping.xlsx describes which column in the source databases has been mapped to which column in BENTH-MARC, whereas taxonomic_filter.xlsx gives an overview of the main taxonomic filter and its validation through WoRMS. File overview File name Size Description Purpose File type occurrences.parquet 210.67 MB All occurrence records in BENTH-MARC Main data package Parquet datasets.csv 972.16 KB All datasets in BENTH-MARC Main data package CSV README.txt 4.46 KB A description of BENTH-MARC Documentation txt workflow.xlsx 14.98 KB A description of the workflow Documentation Excel glossary.xlsx 16.31 KB Column glossary and data dictionary Documentation Excel column_mapping.xlsx 16.47 KB Column mapping matrix from source database to BENTH-MARC Documentation Excel taxonomic_filter.xlsx 55.80 KB Taxonomic filter Documentation Excel map_stations.png 1.37 MB Overview of stations in BENTH-MARC Documentation PNG Aleutian Islands LME.csv 45.84 MB Occurrence records in Aleutian Islands LME One CSV per LME CSV Baffin Bay LME.csv 49.79 MB Occurrence records in Baffin Bay LME One CSV per LME CSV Barents Sea LME.csv 551.00 MB Occurrence records in Barents Sea LME One CSV per LME CSV Beaufort Sea LME.csv 56.31 MB Ocurrence records in Beaufort LME One CSV per LME CSV Central Arctic LME.csv 13.26 MB Occurrence records in Central Arctic LME One CSV per LME CSV Chukchi Sea LME.csv 62.30 MB Occurrence records in Chukchi Sea LME One CSV per LME CSV East Bering Sea LME.csv 33.07 MB Occurrence records in East Bering Sea LME One CSV per LME CSV East Siberian Sea LME.csv 681.34 KB Occurrence records in East Siberian Sea LME One CSV per LME CSV Faroe Plateau LME.csv 3.03 MB Occurrence records in Faroe Plateau LME One CSV per LME CSV Greenland Sea LME.csv 21.06 MB Occurrence records in Greenland Sea LME One CSV per LME CSV Hudson Bay LME.csv 16.43 MB Occurrence records in Hudson Bay LME One CSV per LME CSV Iceland LME.csv 6.17 MB Occurrence records in Iceland LME One CSV per LME CSV Kara Sea LME.csv 9.42 MB Occurrence records in Kara Sea LME One CSV per LME CSV Labrador Sea LME.csv 9.03 MB Occurrence records in Labrador Sea LME One CSV per LME CSV Laptev Sea LME.csv 11.14 MB Occurrence records in Laptev Sea LME One CSV per LME CSV Northern Canadian Archipelago LME.csv 2.70 MB Occurrence records in Northern Canadian Archipelago LME One CSV per LME CSV Norwegian Sea LME.csv 729.75 MB Occurrence records in Norwegian Sea LME One CSV per LME CSV West Bering Sea LME.csv 3.66 MB Occurrence records in West Bering Sea LME One CSV per LME CSV



