EPA-ng: massively parallel evolutionary placement of genetic sequences
收藏资源简介:
Next Generation Sequencing (NGS) technologies have led to a ubiquity of molecular sequence data. This data avalanche is particularly challenging in metagenetics, which focuses on taxonomic identification of sequences obtained from diverse microbial environments. Phylogenetic placement methods determine how these sequences fit into anevolutionary context. Previous implementations of phylogenetic placement algorithms, such as the Evolutionary Placement Algorithm (EPA) included in RAxML, or pplacer, are being increasingly used for this purpose. However, due to the steady progress in NGS technologies, the current implementations face substantial scalability limitations. Here we present EPA-ng, a complete reimplementation of the EPA that is substantially faster, offers a distributed memory parallelization, and integrates concepts from both, RAxML-EPA and pplacer. EPA-ng can be executed on standard shared memory, as well as on distributed memory systems (e.g., computing clusters). To demonstr...
下一代测序(Next Generation Sequencing, NGS)技术已推动分子序列数据实现广泛普及。此类数据洪流在宏遗传学(metagenetics)领域尤为棘手,该领域聚焦于对从多样微生物环境中获取的序列开展分类学鉴定。系统发育放置(phylogenetic placement)方法可用于确定这些序列如何融入进化框架。此前的系统发育放置算法实现,例如RAxML中集成的进化放置算法(Evolutionary Placement Algorithm, EPA)以及pplacer,正日益被应用于该类任务。然而,随着NGS技术的稳步发展,现有算法实现面临显著的可扩展性限制。本文提出EPA-ng,一款对EPA的全新重实现版本,其速度大幅提升,支持分布式内存并行计算,并整合了RAxML-EPA与pplacer的设计理念。EPA-ng可在标准共享内存系统,以及分布式内存系统(例如计算集群)上运行。为验证...




