遇见数据集

Demographic model selection using random forests and the site frequency spectrum

收藏
DataONE2020-06-24 更新2025-05-03 收录
官方服务:

资源简介:

Phylogeographic data sets have grown from tens to thousands of loci in recent years, but extant statistical methods do not take full advantage of these large data sets. For example, approximate Bayesian computation (ABC) is a commonly used method for the explicit comparison of alternate demographic histories, but it is limited by the “curse of dimensionality” and issues related to the simulation and summarization of data when applied to next-generation sequencing (NGS) data sets. We implement here several improvements to overcome these difficulties. We use a Random Forest (RF) classifier for model selection to circumvent the curse of dimensionality and apply a binned representation of the multidimensional site frequency spectrum (mSFS) to address issues related to the simulation and summarization of large SNP data sets. We evaluate the performance of these improvements using simulation and find low overall error rates (~7%). We then apply the approach to data from Haplotrema vancouveren...

近年来,系统地理学数据集(Phylogeographic data sets)的位点数量已从数十个增长至数千个,但现有统计方法尚未充分利用这类大型数据集。例如,近似贝叶斯计算(approximate Bayesian computation, ABC)是用于显式对比不同人口历史模型的常用方法,但在应用于下一代测序(next-generation sequencing, NGS)数据集时,其受限于维度灾难以及与数据模拟和汇总相关的诸多问题。本文在此提出多项改进措施以攻克上述难题:我们采用随机森林(Random Forest, RF)分类器开展模型选择,以规避维度灾难;同时引入多维位点频率谱(multidimensional site frequency spectrum, mSFS)的分箱表示方法,解决大型单核苷酸多态性(Single Nucleotide Polymorphism, SNP)数据集的模拟与汇总相关问题。我们通过模拟实验对这些改进方案的性能进行了评估,结果显示整体错误率较低(约7%)。随后我们将该方法应用于Haplotrema vancouveren……的相关数据。

创建时间:
2025-04-20
二维码
社区交流群
二维码
科研交流群
商业服务