Squeegee: de novo identification of reagent and laboratory induced microbial contaminants in low biomass microbiomes, simulation dataset 0.25% spike-in contaminant sequences
收藏资源简介:
Computational analysis of host-associated microbiomes has opened the door to numerous discoveries relevant to human health and disease. However, contaminant sequences in metagenomic samples can potentially impact the interpretation of findings reported in microbiome studies, especially in low biomass environments. Our hypothesis is that contamination from DNA extraction kits or sampling lab environments will leave taxonomic "bread crumbs” across multiple distinct sample types, allowing for the detection of microbial contaminants when negative controls are unavailable. To test this hypothesis we implemented Squeegee, a de novo contamination detection tool. We tested Squeegee on simulated and real low biomass metagenomic datasets. On the low biomass samples, we compared Squeegee predictions to experimental negative control data and show that Squeegee accurately recovers known contaminants. We also analyzed 749 metagenomic datasets from the Human Microbiome Project and identified likely previously unreported kit contamination. Collectively, our results highlight that Squeegee can identify microbial contaminants with high precision. Simulation Dataset 0.25% contaminant spike-in.
宿主相关微生物组的计算分析,已为诸多与人类健康和疾病相关的科研发现开辟了全新路径。然而,宏基因组样本中的污染序列可能会干扰微生物组研究中已报道结果的解读,该问题在低生物量环境样本中尤为显著。我们提出的假说为:来自DNA提取试剂盒或采样实验室环境的污染,会在多种不同样本类型中留下分类学“足迹”,从而可在缺乏阴性对照的情况下实现微生物污染物的检测。为验证该假说,我们开发了Squeegee——一款从头污染检测工具。我们在模拟数据集与真实低生物量宏基因组数据集上对Squeegee开展了测试。针对低生物量样本,我们将Squeegee的预测结果与实验阴性对照数据进行比对,证实其可准确还原已知污染物。我们还分析了来自人类微生物组计划(Human Microbiome Project)的749个宏基因组数据集,发现了此前可能未被报道的试剂盒污染情况。综上,本研究结果表明,Squeegee可高精度识别微生物污染物。模拟数据集:污染物掺入比例为0.25%。




