Data from: How to optimize the precision of allele and haplotype frequency estimates using pooled-sequencing data
收藏资源简介:
Sequencing pools of individuals rather than individuals separately reduces the costs of estimating allele frequencies at many loci in many populations. Theoretical and empirical studies show that sequencing pools comprising a limited number of individuals (typically fewer than 50) provides reliable allele frequency estimates, provided that the DNA pooling and DNA sequencing steps are carefully controlled. Unequal contributions of different individuals to the DNA pool and the mean and variance in sequencing depth both can affect the standard error of allele frequency estimates. To our knowledge, no study separately investigated the effect of these two factors on allele frequency estimates; so that there is currently no method to a priori estimate the relative importance of unequal individual DNA contributions independently of sequencing depth. We develop a new analytical model for allele frequency estimation that explicitly distinguishes these two effects. Our model shows that the DNA pooling variance in a pooled sequencing experiment depends solely on two factors: the number of individuals within the pool and the coefficient of variation of individual DNA contributions to the pool. We present a new method to experimentally estimate this coefficient of variation when planning a pooled sequencing design where samples are either pooled before or after DNA extraction. Using this analytical and experimental framework, we provide guidelines to optimize the design of pooled sequencing experiments. Finally, we sequence replicated pools of inbred lines of the plant Medicago truncatula and show that the predictions from our model generally hold true when estimating the frequency of known multilocus haplotypes using pooled sequencing.
相较于单独对单个个体进行测序,对多个个体的混合样本池开展测序,可显著降低在多个群体的众多位点上估算等位基因频率(allele frequencies)的成本。理论与实证研究均表明,只要严格控制DNA混样(DNA pooling)与DNA测序的操作步骤,包含有限数量个体(通常少于50个)的混合测序样本池即可提供可靠的等位基因频率估算结果。不同个体对DNA混样的贡献不均,以及测序深度(sequencing depth)的均值与方差,均会影响等位基因频率估算结果的标准误(standard error)。据我们所知,目前尚无研究单独探讨这两种因素对等位基因频率估算的影响,因此当前尚无方法可以在不考虑测序深度的前提下,先验地估算不同个体DNA贡献不均的相对重要性。我们开发了一种全新的分析模型用于等位基因频率估算,该模型能够明确区分上述两种效应。我们的模型显示,混合测序实验中的DNA混样方差仅取决于两个因素:样本池中的个体数量,以及个体对混样的DNA贡献量的变异系数(coefficient of variation)。我们提出了一种全新的实验方法,可在规划混合测序实验设计时估算该变异系数,此时样本可在DNA提取前或DNA提取后进行混合。借助这一分析与实验框架,我们为优化混合测序实验的设计提供了科学指导准则。最后,我们对植物蒺藜苜蓿(Medicago truncatula)的近交系(inbred lines)重复混合样本池进行了测序,结果表明,当使用混合测序技术估算已知多基因座单倍型的频率时,我们模型的预测结果整体与实际观测相符。



