Data from: New approaches for unravelling reassortment pathways
收藏资源简介:
BACKGROUND: Every year the human population encounters epidemic outbreaks of influenza, and history reveals recurring pandemics that have had devastating consequences. The current work focuses on the development of a robust algorithm for detecting influenza strains that have a composite genomic architecture. These influenza subtypes can be generated through a reassortment process, whereby a virus can inherit gene segments from two different types of influenza particles during replication. Reassortant strains are often not immediately recognised by the adaptive immune system of the hosts and hence may be the source of pandemic outbreaks. Owing to their importance in public health and their infectious ability, it is essential to identify reassortant influenza strains in order to understand the evolution of this virus and describe reassortment pathways that may be biased towards particular viral segments. Phylogenetic methods have been used traditionally to identify reassortant viruses. In many studies up to now, the assumption has been that if two phylogenetic trees differ, it is because reassortment has caused them to be different. While phylogenetic incongruence may be caused by real differences in evolutionary history, it can also be the result of phylogenetic error. Therefore, we wish to develop a method for distinguishing between topological inconsistency that is due to confounding effects and topological inconsistency that is due to reassortment. RESULTS: The current work describes the implementation of two approaches for robustly identifying reassortment events. The algorithms rest on the idea of significance of difference between phylogenetic trees or phylogenetic tree sets, and subtree pruning and regrafting operations, which mimic the effect of reassortment on tree topologies. The first method is based on a maximum likelihood (ML) framework (MLreassort) and the second implements a Bayesian approach (Breassort) for reassortment detection. We focus on reassortment events that are found by both methods. We test both methods on a simulated dataset and on a small collection of real viral data isolated in Hong Kong in 1999. CONCLUSIONS: The nature of segmented viral genomes present many challenges with respect to disease. The algorithms developed here can effectively identify reassortment events in small viral datasets and can be applied not only to influenza but also to other segmented viruses. Owing to computational demands of comparing tree topologies, further development in this area is necessary to allow their application to larger datasets.
背景:每年人类都面临流感疫情暴发,历史上也曾出现过多次造成灾难性后果的流感大流行。本研究致力于开发一种鲁棒性算法,用于检测具备复合基因组结构的流感毒株。这类流感亚型可通过重配过程产生:病毒在复制过程中,可从两种不同亚型的流感病毒颗粒中获取基因片段。重配毒株通常不会被宿主的适应性免疫系统即时识别,因此可能成为流感大流行的源头。鉴于其在公共卫生领域的重要性以及自身的传染性,鉴定重配流感毒株对于理解该病毒的演化、解析可能偏向特定病毒片段的重配路径至关重要。传统上,系统发育分析方法被用于鉴定重配病毒。迄今为止的诸多研究均假设:若两棵系统发育树存在差异,则该差异是由重配事件导致的。然而,系统发育不一致性既可能源于真实的演化历史差异,也可能由系统发育分析误差所引发。因此,本研究旨在开发一种方法,以区分由混杂效应导致的拓扑结构不一致,与由重配事件引发的拓扑结构不一致。 结果:本研究实现了两种可稳健识别重配事件的方法。这两种算法均基于系统发育树或系统发育树集之间差异的显著性这一核心思想,并通过子树剪枝与重接(subtree pruning and regrafting, SPR)操作模拟重配对树拓扑结构的影响。第一种方法基于最大似然(maximum likelihood, ML)框架,命名为MLreassort;第二种方法则采用贝叶斯方法,命名为Breassort,用于重配检测。本研究聚焦于两种方法共同识别出的重配事件。我们分别通过模拟数据集,以及1999年在香港分离得到的小型真实病毒数据集对两种方法进行了测试。 结论:分节段病毒基因组的特性给疾病防控带来了诸多挑战。本文开发的算法可在小型病毒数据集内有效识别重配事件,且不仅可应用于流感病毒,还可推广至其他分节段病毒。鉴于比对树拓扑结构的计算复杂度较高,未来仍需进一步优化相关技术,以实现对更大规模数据集的应用。



