遇见数据集

Environmental DNA metabarcoding to monitor tropical reef fishes in Santa Marta

收藏
Figshare2021-06-11 更新2026-04-28 收录
官方服务:

资源简介:

Environmental DNA (eDNA) provides a revolutionary method to monitor species in marine ecosystems from animal DNA present in the water. Examining the capacity of eDNA to provide accurate biodiversity measures in species-rich ecosystems such as coral reefs is a prerequisite for their long-term monitoring. Here, we surveyed a Colombian tropical marine reefs, the Gayraca Bay near Santa Marta using eDNA method. We collected a large quantity of surface water (30 L per filter) above the reefs and applied a metabarcoding protocol using three different primer sets targeting the 12S mitochondrial DNA, specific to vertebrates, Actinopterygii and Elasmobranchii. The assignment of eDNA sequences to species using a public reference database allowed detecting the presence of 85 fish species, 92 genera and 57 families in Providencia. Filtering and taxonomic assignments Obitools clustering: Following the sequencing, reads were processed to remove errors and analyzed using programs implemented in the OBITools package (http://metabarcoding.org/obitools; Boyer et al., 2016) following a previous protocol (Valentini et al., 2016). We assembled the forward and reverse reads using the ILLUMINAPAIREDEND program using a minimum score of 40 and retrieving only joined sequences. Then, we assigned the reads to each sample using NGSFILTER software. A separate data set was created for each sample by splitting the original data set into several files using OBISPLIT. After this step, we analyzed each data set sample individually before merging the taxon list for the final ecological analysis. Strictly identical sequences were clustered together using OBIUNIQ. We excluded sequences shorter than 20 bp or with fewer than 10 reads using the OBIGREP program and ran the OBICLEAN program within a PCR product. All sequences labeled ‘internal’ that most likely corresponded to PCR substitutions and indel errors were discarded. We realized the taxonomic assignment of the remaining sequences using the program ECOTAG using the NCBI reference database (www.ncbi.nlm.nih.gov, release 233, downloaded on 11 Oct. 2019). We corrected taxonomic assignment outputs to avoid any over-confidence in assignments: species-level assignments were validated only for sequences with an identification match >98%, genus-level for a 96-98% match and family-level for an 90-96% match. Considering the wrong assignment of a few sequences to the sample due to tag-jumps (Schnell et al., 2015), we discarded all sequences with a frequency of occurrence Swarm clustering: We applied a second bioinformatics workflow, the clustering algorithm SWARM, which uses sequence similarity and abundance patterns to cluster multiple variants of sequences into MOTU (Molecular Operational Taxonomic Units; Mahé et al., 2014; Rognes et al., 2016). While the OBITools bioinformatics pipeline allows optimizing the taxonomic identification of sequence, even or rare ones, the SWARM approach allows clustering similar sequence and provides full compositional matrices even in the absence of a complete reference database (Marques et al., 2020). First, we merged sequences using vsearch software to remove sequences containing ambiguities (Rognes et al., 2016). We then applied CUTADAPT software (Martin, 2013) for demultiplexing and primer trimming (Table TS2). Next, we ran SWARM with a minimum distance of one mismatch to make clusters. Once the MOTUs were generated, we used the most abundant sequence within each cluster as a representative sequence for taxonomic assignment. Then, we applied a post-clustering curation algorithm (LULU; Frøslev et al., 2017) to curate the data. We validated the outputs using the same thresholds as for the OBITools one. Further quality cleaning was identical to that used in the OBITools pipeline (identify minimum number of reads, remove non-target taxa, apply tag-jump cleaning), with the addition of a single step removing all MOTUs present in only PCR within the entire data set. This additional step was necessary because PCR errors are unlikely to be present in more than one PCR occurrence, and it removes spurious MOTUs that would otherwise inflate diversity estimates (see Marques et al., 2020). For the teleo marker, this approach has been validated with fish observation data, where MOTUs generally correspond to species (Marques et al., 2020), but estimates were so far not validated for other markers.

环境DNA(eDNA)是一种革命性的方法,可通过水体中留存的动物DNA监测海洋生态系统中的物种。在物种丰富的生态系统(如珊瑚礁)中,验证eDNA提供准确生物多样性评估数据的能力,是开展其长期监测的前提条件。本研究针对哥伦比亚圣玛尔塔附近的盖拉卡湾热带海洋珊瑚礁,采用eDNA技术开展调查。我们采集了珊瑚礁上方的大量表层水样(每滤膜过滤30 L),并采用针对脊椎动物、辐鳍鱼纲(Actinopterygii)和软骨鱼纲(Elasmobranchii)的三套不同引物组,基于12S线粒体DNA建立宏条形码(metabarcoding)实验流程。通过公共参考数据库对eDNA序列进行物种注释,在普罗维登西亚海域共检测到85种鱼类、92个属及57个科。 序列过滤与分类注释(OBITools聚类):测序完成后,我们对测序读段(reads)进行错误去除处理,并参考既往实验方案(Valentini等,2016),采用OBITools工具包(http://metabarcoding.org/obitools; Boyer等,2016)中的程序开展分析。首先使用ILLUMINAPAIREDEND程序拼接正向与反向测序读段(reads),设置最低得分阈值为40,仅保留拼接成功的序列。随后通过NGSFILTER软件将测序读段(reads)分配至对应样本,利用OBISPLIT工具将原始数据集拆分为多个样本专属文件。完成此步骤后,我们先单独分析每个样本的数据集,再合并分类单元列表用于最终的生态分析。使用OBIUNIQ将完全一致的序列聚类为一组,通过OBIGREP程序剔除长度短于20 bp或测序读段(reads)数少于10的序列,并在PCR产物范围内运行OBICLEAN程序,丢弃所有标记为“internal”、大概率对应PCR替换及插入缺失错误的序列。随后使用ECOTAG程序,结合美国国家生物技术信息中心(NCBI)参考数据库(www.ncbi.nlm.nih.gov,版本233,2019年10月11日下载)对剩余序列进行物种注释。为避免注释结果过度置信,我们对注释结果进行了校正:物种水平注释仅匹配度>98%的序列有效,属水平注释匹配度为96%~98%,科水平注释匹配度为90%~96%。考虑到标签跳变(tag-jump)可能导致少量序列被错误分配至样本(Schnell等,2015),我们剔除了出现频率相关序列,随后开展SWARM聚类分析。 SWARM聚类流程:我们采用第二种生物信息学分析流程——聚类算法SWARM,该算法利用序列相似性与丰度模式,将序列的多种变异聚类为分子操作分类单元(MOTU, Molecular Operational Taxonomic Units; Mahé等,2014; Rognes等,2016)。相较于OBITools分析流程可优化序列的分类鉴定能力(包括稀有序列),SWARM方法可对相似序列进行聚类,且即使缺乏完整的参考数据库,也可生成完整的组成矩阵(Marques等,2020)。首先,我们使用vsearch软件合并序列,以去除含歧义碱基的序列(Rognes等,2016)。随后使用CUTADAPT软件(Martin,2013)完成样本拆分与引物切除(详见附表TS2)。接下来,设置最小错配数为1,运行SWARM进行聚类。生成MOTU后,我们选取每个聚类中丰度最高的序列作为代表序列用于物种注释。随后应用聚类后校正算法(LULU; Frøslev等,2017)对数据进行校正。我们采用与OBITools流程相同的阈值验证注释结果。后续的质量清洗步骤与OBITools流程一致(确定最小测序读段(reads)数、剔除非目标分类单元、进行标签跳变清洗),额外增加了一步剔除仅在一次PCR反应中出现的所有MOTU的步骤。该额外步骤十分必要,因为PCR错误不太可能在多次PCR反应中同时出现,可去除那些会虚增多样性评估结果的虚假MOTU(详见Marques等,2020)。针对teleo标记(teleo marker),该方法已通过鱼类观测数据得到验证,其MOTU通常可对应至物种水平(Marques等,2020),但目前尚未针对其他标记物开展相关验证。

创建时间:
2021-06-11
二维码
社区交流群
二维码
科研交流群
商业服务