Script and data from: The best of two worlds: toward large-scale monitoring of biodiversity combining metabarcoding and optimised parataxonomic validation.
收藏资源简介:
Description Publication abstract 1- In a context of unprecedented biodiversity decline, there is a critical need for reliable monitoring tools to measure species diversity and their dynamic at large scales. DNA-based identification methods, including metabarcoding, were proposed as an effective way to reach this aim. However, these molecular methods suffer from intrinsic limitations (e.g. false positive and false negative species identifications) that are difficult to quantify in the field, and do not allow the estimation of species abundance. 2- To overcome these obstacles, we propose the Human-Assisted Molecular Identification (HAMI) framework, a semi-automated method based on a combination of metabarcoding and image-based parataxonomic validation of outputs and recording of abundance data. This approach was tested on beetles (Coleoptera) within the 500-ENI network, a national biodiversity monitoring initiative covering more than 500 agricultural field margins in mainland France. We assessed the advantages of using HAMI over the exclusive use of a metabarcoding approach by examining 491 samples. 3- We show that an average of 23% of the specific composition is missed when relying exclusively on metabarcoding, this percent being consistently higher in the most species-rich samples. These false negatives are the result of unavoidable primer bias and stochasticity in tissue quality. In addition, on average 20% of the species composition identified by molecular-only approaches correspond to false positives linked to cross-sample contaminations or mis-identified barcode sequences in databases. 4- The combination of molecular methodologies and parataxonomic validation in HAMI significantly reduces metabarcoding’s intrinsic bias and recovers abundance data, thus providing appropriate data for large scale monitoring of biodiversity. This approach also streamlines the time-consuming parataxonomist expertise, enabling users to engage in a virtuous circle of database improvement through the identification of specimens associated with missing barcodes, and identification of barcodes with erroneous assignations. As such, HAMI fills an important gap in large biodiversity monitoring, providing for a better equilibrium between the power of molecular approaches and the need of human expertise. File description: MiSeq raw sequences of the COI barcode from 491 Coleoptera field samples : The Raw_sequencage_data ZIP directory contains the FASTQ files of the paired-end reads (R1: reads 1; R2: reads 2) produced for each Coleoptera field samples in duplicate using the MiSeq platform GenSeq (ISEM - University of Montpellier) The HAMI_data_script_results_R zip directory contains Rmarkdown script files (.Rmd and .html) and associated data used to analyse the systemic errors of the metabarcoding approach (N= 491 Coleoptera field samples). The HAMI_pipeline zip directory contains all the codes associated with the HAMI pipeline, as well as a ReadMe file and a test data set.



