Data from: Demographic model selection using random forests and the site frequency spectrum
收藏资源简介:
Phylogeographic data sets have grown from tens to thousands of loci in recent years, but extant statistical methods do not take full advantage of these large data sets. For example, approximate Bayesian computation (ABC) is a commonly used method for the explicit comparison of alternate demographic histories, but it is limited by the “curse of dimensionality” and issues related to the simulation and summarization of data when applied to next-generation sequencing (NGS) data sets. We implement here several improvements to overcome these difficulties. We use a Random Forest (RF) classifier for model selection to circumvent the curse of dimensionality and apply a binned representation of the multidimensional site frequency spectrum (mSFS) to address issues related to the simulation and summarization of large SNP data sets. We evaluate the performance of these improvements using simulation and find low overall error rates (~7%). We then apply the approach to data from Haplotrema vancouverense, a land snail endemic to the Pacific Northwest of North America. Fifteen demographic models were compared, and our results support a model of recent dispersal from coastal to inland rainforests. Our results demonstrate that binning is an effective strategy for the construction of a mSFS and imply that the statistical power of RF when applied to demographic model selection is at least comparable to traditional ABC algorithms. Importantly, by combining these strategies, large sets of models with differing numbers of populations can be evaluated.
近年来,系统发育地理数据集的基因座数量已从数十个增长至数千个,但现存的统计方法未能充分挖掘这类大规模数据集的潜力。例如,近似贝叶斯计算(approximate Bayesian computation,ABC)是用于明确比较不同种群历史动态的常用方法,但在应用于下一代测序(next-generation sequencing,NGS)数据集时,其受限于“维度灾难”,且存在数据模拟与汇总相关的问题。本研究针对上述难点提出了多项改进策略:采用随机森林(Random Forest,RF)分类器进行模型选择以规避维度灾难,同时对多维位点频率谱(multidimensional site frequency spectrum,mSFS)采用分箱表征,以解决大规模单核苷酸多态性(Single Nucleotide Polymorphism,SNP)数据集的模拟与汇总难题。本研究通过模拟实验对这些改进策略的性能进行评估,结果显示整体错误率较低(约7%)。随后,我们将该方法应用于北美太平洋西北沿岸特有陆生蜗牛Haplotrema vancouverense的测序数据。研究共比较了15种种群历史模型,结果支持“从沿海雨林向内陆雨林发生近期扩散”的模型。本研究结果表明,分箱处理是构建多维位点频率谱的有效策略,同时证实,将随机森林分类器应用于种群历史模型选择时,其统计效力至少可与传统近似贝叶斯计算算法相媲美。尤为重要的是,通过整合上述策略,即可对包含不同种群数量的大规模模型集开展评估。



