Pangenome Mutation-Annotated Networks
收藏资源简介:
Pangenomics is an emerging field that uses a collection of genomes of a species instead of a single reference genome to overcome reference bias and study the within-species genetic diversity. Future pangenomics applications will require analyzing large and ever-growing collections of genomes. Therefore, the choice of data representation is a key determinant of the scope, as well as the computational and memory performance of pangenomic analyses. Current pangenome formats, while capable of storing genetic variations across multiple genomes, fail to capture the shared evolutionary and mutational histories among them, thereby limiting their applications. They are also inefficient for storage, and therefore face significant scaling challenges. In this manuscript, we propose PanMAN, a novel data structure that is information-wise richer than all existing pangenome formats – in addition to representing the alignment and genetic variation in a collection of genomes, PanMAN represents the shared mutational and evolutionary histories inferred between those genomes. By using “evolutionary compression”, PanMAN achieves 5.2 to 680-fold compression over other variation-preserving pangenomic formats. PanMAN's relative performance generally improves with larger datasets and it is compatible with any method for inferring phylogenies and ancestral nucleotide states. Using SARS-CoV-2 as a case study, we show that PanMAN offers a detailed and accurate portrayal of the pathogen's evolutionary and mutational history, facilitating the discovery of new biological insights. We also present panmanUtils, a software toolkit that supports common pangenomic analyses and makes PanMANs interoperable with existing tools and formats. PanMANs are poised to enhance the scale, speed, resolution, and overall scope of pangenomic analyses and data sharing.
泛基因组学(Pangenomics)是一门新兴研究领域,其摒弃单一参考基因组,转而采用某一物种的全基因组集合进行研究,以克服参考基因组偏倚问题,并探究物种内的遗传多样性。未来泛基因组学的应用将需要处理规模庞大且持续增长的基因组集合,因此数据表示方式的选择是决定泛基因组分析的研究范围、计算性能与内存占用的关键因素。当前的泛基因组格式虽可存储多基因组间的遗传变异,但无法捕捉这些基因组共有的进化与突变历史,从而限制了其应用;同时这类格式的存储效率低下,因此面临显著的规模化挑战。本文提出了PanMAN这一新型数据结构,其信息丰富度优于所有现有泛基因组格式:除可表征基因组集合中的序列比对与遗传变异外,PanMAN还能呈现这些基因组间推导得到的共有的突变与进化历史。通过运用“进化压缩”技术,PanMAN相较于其他保留变异的泛基因组格式,可实现5.2倍至680倍的压缩比。PanMAN的相对性能通常随数据集规模扩大而提升,且兼容所有用于推导系统发育与祖先核苷酸状态的分析方法。以严重急性呼吸综合征冠状病毒2(SARS-CoV-2)作为案例研究对象,我们证明PanMAN可精准且细致地刻画该病原体的进化与突变历史,助力发掘全新的生物学见解。我们还推出了panmanUtils软件工具包,可支持常见的泛基因组分析操作,并实现PanMAN与现有工具及格式的互操作性。PanMAN有望大幅提升泛基因组分析与数据共享的规模、速度、分辨率及整体研究范围。



