PCR.PTML-Project Supporting Information
收藏资源简介:
The characterization of DNA sequence variants (variant calling) from ancient samples (aDNA) is a topic of major interest in modern science. However, there are important theoretical and experimental problems to be overcome like the low number of reliable samples, the possibility of contamination, <i>etc</i>. On the other hand, the main concern during the extraction ancient of DNA (aDNA) is to prevent contamination with modern materials. Consequently, it is crucial to be able to differentiate modern from ancient sequences in aDNA variant calling. In this context, computational techniques may play an important role. Most methods reported use alignment-dependent algorithms relying on limited samples. More recently Artificial Intelligence / Machine Learning (AI/ML) methods have been proposed as an alternative. However, almost all of them focus on the analysis of Mammoth, Neanderthal, or other Demographic aDNA problems and omit the study of aDNA paleo-microobiomes. In this work, we report by the first time the PCR.PTML methodology based on the combination of experimental Polymerase Chain Reaction (PCR) metagenomic analysis of aDNA sequences with the Perturbation Theory (PT) and Machine Learning (ML) predictive modeling. Firstly, we reported the extraction and PCR analysis of a new set of putative microbiome 16S aDNA sequences from Miocene fossil amber. Next, we developed a variant calling PTML model able to discriminate 16S aDNA sequences of Miocene bacteria from modern bacteria sequences with more that 80% of specificity and sensitivity in training and validation series. We used both Linear Discriminant Analysis (LDA) and Artificial Neural Networks (ANN) algorithms for variant calling exploration of 100000 combinations of query and reference sequences to seek model.
从古代样本(古代DNA,aDNA)中开展DNA序列变异表征(变异识别,variant calling)是现代科学领域的重要研究课题。然而该方向仍存在诸多亟待解决的理论与实验难题,例如可靠样本量稀缺、存在外源污染风险等。另一方面,古代DNA(aDNA)提取过程中的核心关切之一是避免现代外源物质污染,因此在aDNA变异识别中实现现代序列与古代序列的精准区分至关重要。在此背景下,计算技术可发挥关键作用。目前已报道的多数方法采用依赖序列比对的算法,且依赖有限的样本规模。近年来,人工智能/机器学习(AI/ML)方法被提出作为替代方案,但此类方法几乎均聚焦于猛犸象、尼安德特人或其他人群相关的古代DNA研究,而忽略了古代微生物组(paleo-microbiomes)的相关分析。本研究首次报道了PCR.PTML方法学,该方法将古代DNA序列的实验聚合酶链式反应(Polymerase Chain Reaction, PCR)宏基因组分析、扰动理论(Perturbation Theory, PT)与机器学习(Machine Learning, ML)预测建模相结合。首先,我们从中新世化石琥珀中提取并PCR分析了一批全新的推定微生物组16S aDNA序列;随后,我们开发了一款变异识别PTML模型,可实现中新世细菌16S aDNA序列与现代细菌序列的精准区分,在训练集与验证集上的特异性与灵敏度均超过80%。本研究采用线性判别分析(Linear Discriminant Analysis, LDA)与人工神经网络(Artificial Neural Networks, ANN)两种算法,对100000条查询序列与参考序列的组合开展变异识别探索以筛选最优模型。



