Data from: An efficient independence sampler for updating branches in Bayesian Markov chain Monte Carlo sampling of phylogenetic trees
收藏资源简介:
Sampling tree space is the most challenging aspect of Bayesian phylogenetic inference. The sheer number of alternative topologies is problematic by itself. In addition, the complex dependency between branch lengths and topology increases the difficulty of moving efficiently among topologies. Current tree proposals are fast but sample new trees using primitive transformations or re-mappings of old branch lengths. This reduces acceptance rates and presumably slows down convergence and mixing. Here, we explore branch proposals that do not rely on old branch lengths but instead are based on approximations of the conditional posterior. Using a diverse set of empirical data sets, we show that most conditional branch posteriors can be accurately approximated via a Γ distribution. We empirically determine the relationship between the logarithmic conditional posterior density, its derivatives, and the characteristics of the branch posterior. We use these relationships to derive an independence sampler for proposing branches with an acceptance ratio of ∼90% on most data sets. This proposal samples branches between 2× and 3× more efficiently than traditional proposals with respect to the effective sample size per unit of runtime. We also compare the performance of standard topology proposals with hybrid proposals that use the new independence sampler to update those branches that are most affected by the topological change. Our results show that hybrid proposals can sometimes noticeably decrease the number of generations necessary for topological convergence. Inconsistent performance gains indicate that branch updates are not the limiting factor in improving topological convergence for the currently employed set of proposals. However, our independence sampler might be essential for the construction of novel tree proposals that apply more radical topology changes.
树空间采样是贝叶斯系统发育推断(Bayesian phylogenetic inference)中最具挑战性的环节。仅备选拓扑结构(topologies)的庞大数量本身就带来了诸多难题。此外,分支长度(branch lengths)与拓扑结构间的复杂依赖关系,进一步提升了在不同拓扑间高效遍历的难度。现有的树提议算法(tree proposals)虽具备较快的运行速度,但均通过对旧分支长度的原始变换或重映射来生成新树,这会降低接受率(acceptance rates),并大概率延缓收敛与混合效率。本文针对不依赖旧分支长度、转而基于条件后验(conditional posterior)近似的分支提议方法展开探索。借助多组多样化的实证数据集(empirical data sets),我们证实绝大多数条件分支后验均可通过伽马(Γ)分布实现精准近似。我们通过实证明确了对数条件后验密度(logarithmic conditional posterior density)、其导数与分支后验特征之间的关联。依托这些关联,我们推导出一款独立采样器(independence sampler),该采样器在绝大多数数据集上可实现约90%的分支接受率。以单位运行时间的有效样本量(effective sample size)为评价标准,该提议算法的分支采样效率较传统方法提升2至3倍。我们还将标准拓扑提议算法与混合提议算法(hybrid proposals)进行了对比,后者借助新型独立采样器更新受拓扑变化影响最大的分支。结果表明,混合提议算法有时可显著减少拓扑收敛所需的世代数。但性能提升存在不一致性,这说明在当前使用的提议算法框架下,分支更新并非提升拓扑收敛性能的瓶颈。不过,我们提出的独立采样器或许可为构建采用更激进拓扑变换的新型树提议算法提供必要支撑。



