遇见数据集

Assessing the Consequences of Denoising Marker-Based Metagenomic Data

收藏
Figshare2016-01-18 更新2026-04-29 收录
官方服务:

资源简介:

Early marker-based metagenomic studies were performed without properly accounting for the effects of noise (sequencing errors, PCR single-base errors, and PCR chimeras). Denoising algorithms have been developed, but they were validated using data derived from mock communities, in which the true sequences were known. Since the algorithms were designed to be used in real community studies, it is important to evaluate the results in such cases. With this goal in mind, we processed a real 16S rRNA metagenomic dataset through five denoising pipelines. By reconstituting the sequence reads at each stage of the pipelines, we determined how the reads were being altered. In one denoising pipeline, AmpliconNoise, we found that the algorithm that was designed to remove pyrosequencing errors changed the reads in a manner inconsistent with the known spectrum of these errors, until one of the parameters was increased substantially from its default value. Additionally, because the longest read was picked as the representative for each cluster, sequences were added to the 3′ ends of shorter reads that were often dissimilar from what had been removed by the truncations of the previous filtering step. In QIIME, the denoising algorithm caused a much larger number of changes to the reads unless the parameters were changed from their defaults. The denoising pipeline in mothur avoided some of these negative side-effects because of its strict default filtering criteria, but these criteria also greatly limited the sequence information produced at the end of the pipeline. We recommend that those using these denoising pipelines be cognizant of these issues and examine how their reads are being transformed by the denoising process as a component of their analysis.

早期基于标记基因的宏基因组学研究,均未妥善考量噪声(测序错误、PCR单碱基错误及PCR嵌合体)所带来的影响。目前已开发出多种去噪算法,但相关算法的验证均基于模拟群落(mock community)测序数据——此类数据的真实序列信息已知。鉴于此类算法的应用场景为真实群落研究,因此在实际样本场景中评估其效果至关重要。基于此研究目标,我们采用五种去噪流程对一套真实16S核糖体RNA(16S rRNA)宏基因组数据集进行了处理。通过在各流程节点重构序列读段(sequence reads),我们明确了读段的具体变化方式。在AmpliconNoise这一去噪流程中,我们发现:本应去除焦磷酸测序错误的算法,其对读段的修改方式与已知的该类错误特征谱并不相符,直至将某一参数从默认值大幅调高后,该匹配问题才得以纠正。此外,由于该流程选取最长读段作为每个聚类的代表序列,因此会将序列添加至较短读段的3'端,而新增序列往往与前一步过滤截断所移除的序列特征并不相似。在QIIME的去噪算法中,除非将参数从默认值调整,否则算法会对读段产生远多于预期的修改。而mothur的去噪流程由于采用了严格的默认过滤标准,因此规避了部分上述负面效应,但该标准也大幅限制了流程最终输出的序列信息量。我们建议,使用此类去噪流程的研究者应充分知晓上述问题,并将检视读段在去噪过程中的变化作为分析流程的必要组成部分。

创建时间:
2016-01-18
二维码
社区交流群
二维码
科研交流群
商业服务