Data from: Sequence clustering threshold has little effect on the recovery of microbial community structure
收藏资源简介:
Analysis of microbial community structure by multivariate ordination methods, using data obtained by high throughput sequencing of amplified markers (i.e., DNA metabarcoding), often requires clustering of DNA sequences into operational taxonomic units (OTUs). Parameters for the clustering procedure tend not to be justified but are set by tradition rather than being based on explicit knowledge. In this study, we explore the extent to which ordination results are affected by variation in parameter settings for the clustering procedure. Amplicon sequence data from nine microbial community studies, representing different sampling designs, spatial scales and ecosystems, were subjected to clustering into OTUs at seven different similarity thresholds (clustering thresholds) ranging from 87% to 99% sequence similarity. The 63 data sets thus obtained were subjected to parallel DCA and GNMDS ordinations. The resulting community structures were highly similar across all clustering thresholds. We explain this pattern by the existence of strong ecological structuring gradients and phylogenetically diverse sets of abundant OTUs that are highly stable across clustering thresholds. Removing low abundance, rare OTUs had negligible effects on community patterns. Our results indicate that microbial data sets with a clear gradient structure are highly robust to choice of sequence clustering threshold.
采用多元排序方法分析微生物群落结构时,通常需基于扩增标记高通量测序所得数据(即DNA宏条形码(DNA metabarcoding))将DNA序列聚类为操作分类单元(OTUs)。聚类流程的参数往往缺乏合理依据,多基于传统经验而非明确的科学知识设定。本研究探讨了聚类流程参数设置的变化对排序结果的影响程度。我们选取9项涵盖不同采样设计、空间尺度与生态系统的微生物群落研究的扩增子序列数据,以87%至99%的序列相似性为区间,设置7种不同相似性阈值(聚类阈值),对序列进行OTU聚类。对由此得到的63个数据集分别并行开展去趋势对应分析(Detrended Correspondence Analysis, DCA)与广义非度量多维尺度分析(Generalized Non-metric Multidimensional Scaling, GNMDS)。所有聚类阈值下得到的群落结构均高度相似。我们将该现象归因于存在较强的生态结构梯度,以及在各聚类阈值下均保持高度稳定的、系统发育多样性丰富的优势OTUs。去除低丰度稀有OTUs对群落格局的影响可忽略不计。本研究结果表明,具备清晰梯度结构的微生物数据集,对序列聚类阈值的选择具有极强的稳健性。



