遇见数据集

Dataset for "To denoise or to cluster? That is not the question. Optimizing pipelines for COI metabarcoding and metaphylogeography

收藏
Mendeley Data2021-01-06 更新2026-04-09 收录
官方服务:

资源简介:

This dataset contains the relevant files for a study optimizing and combining denoising and clustering algorithms for COI metabarcoding. The abstract is: The recent blooming of metabarcoding applications to biodiversity studies comes with some relevant methodological debates. One such issue concerns the treatment of reads by denoising or by clustering methods, which have been wrongly presented as alternatives. It has also been suggested that denoised sequence variants should replace clusters as the basic unit of metabarcoding analyses, missing the fact that sequence clusters are a proxy for species-level entities, the basic unit in biodiversity studies. We argue here that methods developed and tested for ribosomal markers have been uncritically applied to highly variable markers such as cytochrome oxidase I (COI) without conceptual or operational (e.g., parameter setting) adjustment. COI has a naturally high intraspecies variability that should be assessed and reported, as it is a source of highly valuable information. We contend that denoising and clustering are not alternatives. Rather, they are complementary and both should be used together in COI metabarcoding pipelines. Using a typical dataset from benthic marine communities, we compared two denoising procedures (based on the UNOISE3 and the DADA2 algorithms), set suitable parameters for denoising and clustering COI datasets, and compared the outcome of applying these processes in different orders. Our results indicate that denoising based on the UNOISE3 algorithm preserves a higher intra-cluster variability. We suggest and test ways to improve this algorithm taking into account the natural variability of each codon position in coding genes. The order of the steps has little influence on the final outcome. We recommend researchers to consider reporting their results in terms of both denoised sequences (a proxy for haplotypes) and clusters formed (a proxy for species), and to avoid collapsing the sequences of the latter into a single representative. This will allow studies at the cluster (ideally equating species-level diversity) and at the intra-cluster level, and will ease additivity and comparability between studies.

本数据集包含一项针对细胞色素氧化酶I (cytochrome oxidase I,简称COI) 元条形码 (metabarcoding) 技术中降噪与聚类算法优化及组合的研究相关文件。该研究的摘要如下:近年来元条形码技术在生物多样性研究中的应用蓬勃发展,但也伴随了若干相关的方法学争议。其中一项争议围绕测序读段的处理方式展开:降噪方法与聚类方法曾被错误地视为互斥替代方案。另有观点提出,经降噪处理的序列变异体应取代聚类单元,成为元条形码分析的基本单位,但该观点忽略了一个核心事实:序列聚类单元是物种级实体的替代表征,而物种级实体正是生物多样性研究的基本分析单位。本研究指出,针对核糖体标记基因开发并验证的分析方法,未经概念或操作层面(如参数设定)的调整,就被盲目套用至COI这类高变异性标记基因的分析中。COI基因天然存在较高的种内变异,而这一特性可作为极具价值的研究信息来源,理应得到评估与报告。我们认为降噪与聚类并非互斥方案,二者具有互补性,应在COI元条形码分析流程中结合使用。本研究利用一套典型的底栖海洋群落数据集,对比了两种基于UNOISE3与DADA2算法的降噪流程,为COI数据集的降噪与聚类步骤设定了适宜参数,并比较了不同执行顺序下的分析结果。研究结果显示,基于UNOISE3算法的降噪流程可保留更高的聚类内变异水平。我们提出并验证了一种改进该算法的方法,该方法充分考虑了编码基因中每个密码子位点的天然变异特性。分析步骤的执行顺序对最终结果影响极小。我们建议研究者在报告结果时,同时呈现经降噪的序列(作为单倍型的替代表征)与所形成的聚类单元(作为物种的替代表征),并避免将后者的序列压缩为单一代表序列。此举既支持针对聚类单元(理想情况下可对应物种水平多样性)与聚类内水平的研究,也能提升不同研究间的可加性与可比性。

创建时间:
2021-01-06
二维码
社区交流群
二维码
科研交流群
商业服务