遇见数据集

test2

收藏
Zenodo2024-12-19 更新2026-05-26 收录
官方服务:

资源简介:

Publication abstract In a context of unprecedented biodiversity decline, there is a critical need for reliable monitoring tools to measure species diversity and their dynamic at large scales. DNA-based identification methods, e.g. metabarcoding, were proposed as an effective way to reach this aim. However, metabarcoding suffers from intrinsic limitations (e.g. false positive and false negative species identifications) that are difficult to quantify in the field, and do not allow the estimation of species abundance. To overcome these obstacles, we propose the Human-Assisted Molecular Identification (HAMI) framework, based on a combination of metabarcoding and image-based parataxonomic validation of outputs and recording of abundance data. This approach was tested on beetles from a national biodiversity monitoring initiative. We assessed the advantages of using HAMI over the exclusive use of a metabarcoding approach by examining 491 samples. We show that an average of 23% of the specific composition is missed when relying exclusively on metabarcoding, this percent being consistently higher in the most species-rich samples. In addition, on average 20% of the species composition identified by molecular-only approaches correspond to false positives linked to cross-sample contaminations or mis-identified barcode sequences in databases. The combination of molecular methodologies and parataxonomic validation in HAMI significantly reduces metabarcoding’s intrinsic bias and recovers reliable abundance data. This approach also enabling users to engage in a virtuous circle of database improvement through the identification of specimens associated with missing barcodes or barcodes with erroneous assignations. As such, HAMI fills an important gap in large-scale biodiversity monitoring by providing appropriate data. File description: MiSeq raw sequences of the COI barcode from 491 Coleoptera field samples : The Raw_sequencage_data ZIP directory contains the FASTQ files of the paired-end reads (R1: reads 1; R2: reads 2) produced for each Coleoptera field samples in duplicate using the MiSeq platform GenSeq (ISEM - University of Montpellier) The HAMI_data_script_results_R zip directory contains Rmarkdown script files (.Rmd and .html) and associated data used to analyse the systemic errors of the metabarcoding approach (N= 491 Coleoptera field samples). The HAMI_pipeline zip directory contains all the codes associated with the HAMI pipeline, as well as a ReadMe file and a test data set. The Residual_chimera.zip directory contains lists of MOTUs associated to residual chimeric sequences that were not filtered using FROGS pipeline but secondarily detected with the de novo approach implemented in HAMI pipeline with ‘isBimeraDenovo’ R function from DADA2 v1.28.0. It contains two distinct files according to the two sequencing runs. The NUMTS_filtered.zip directory contains lists of MOTUs that were excluded of the final dataset according to the NUMTS filtering. File xxx_pseudogene_f1_deteled.csv corresponds to MOTUs that were excluded according to the first filtrering step based on DNA sequencing. File xxx_pseudogene_f2_deteled.csv corresponds to merged MOTUs that were excluded according to the second filter based on occurrence and percentage of identity. This folder contains files for the two sequencing runs.

提供机构:
Zenodo
创建时间:
2024-04-15
二维码
社区交流群
二维码
科研交流群
商业服务