Unforeseen consequences of excluding missing data from next-generation sequences: simulation study of RAD sequences
收藏资源简介:
There is a lack of consensus on how next-generation sequence data should be considered for phylogenetic and phylogeographic estimates, with some studies excluding loci with missing data, while others include them, even when sequences are missing from a large number of individuals. Here we use simulations, focusing specifically on RAD sequences, to highlight some of the unforeseen consequence of excluding missing data from next-generation sequencing. Specifically, we show that in addition to the obvious effects associated with reducing the amount of data used to make historical inferences, the decisions we make about missing data (such as the minimum number of individuals with a sequence for a locus to be included in the study) also impact the types of loci sampled for a study. In particular, as the tolerance for missing data becomes more stringent, the mutational spectrum represented in the sampled loci becomes truncated such that loci with the highest mutation rates are disproportionat...
当前学界对于如何在系统发育与系统地理学推断中应用下一代测序数据(next-generation sequence data)尚未达成共识:部分研究会剔除存在缺失数据的基因座(locus,复数loci),而另一些研究则保留此类基因座,即便某基因座在大量个体中均存在序列缺失的情况。本研究通过模拟实验,且重点聚焦于RAD序列(RAD sequences),旨在阐明从下一代测序数据中剔除缺失数据可能带来的若干未预见后果。具体而言,本研究表明:除了因缩减用于历史演化推断的数据体量所带来的显而易见的影响外,我们针对缺失数据制定的处理规则(例如:某基因座纳入研究所需的、具备有效序列的个体最小数量),也会对研究中采样的基因座类型造成影响。尤为值得注意的是,当对缺失数据的容忍阈值愈发严格时,采样基因座所涵盖的突变谱(mutational spectrum)会发生截断,使得突变率最高的基因座以不成比例的



