Data from: Genotyping-by-sequencing for Populus population genomics: an assessment of genome sampling patterns and filtering approaches
收藏资源简介:
Continuing advances in nucleotide sequencing technology are inspiring a suite of genomic approaches in studies of natural populations. Researchers are faced with data management and analytical scales that are increasing by orders of magnitude. With such dramatic advances comes a need to understand biases and error rates, which can be propagated and magnified in large-scale data acquisition and processing. Here we assess genomic sampling biases and the effects of various population-level data filtering strategies in a genotyping-by-sequencing (GBS) protocol. We focus on data from two species of Populus, because this genus has a relatively small genome and is emerging as a target for population genomic studies. We estimate the proportions and patterns of genomic sampling by examining the Populus trichocarpa genome (Nisqually-1), and demonstrate a pronounced bias towards coding regions when using the methylation-sensitive ApeKI restriction enzyme in this species. Using population-level data from a closely related species (P. tremuloides), we also investigate various approaches for filtering GBS data to retain high-depth, informative SNPs that can be used for population genetic analyses. We find a data filter that includes the designation of ambiguous alleles resulted in metrics of population structure and Hardy-Weinberg equilibrium that were most consistent with previous studies of the same populations based on other genetic markers. Analyses of the filtered data (27,910 SNPs) also resulted in patterns of heterozygosity and population structure similar to a previous study using microsatellites. Our application demonstrates that technically and analytically simple approaches can readily be developed for population genomics of natural populations.
核苷酸测序技术的持续迭代,正为自然种群研究领域催生一系列基因组学研究手段。当前研究人员面临的数据管理与分析规模正呈数量级增长,伴随着这类技术的迅猛发展,亟需明确各类偏差与错误率——这类偏差与错误会在大规模数据获取与处理环节中被传递并放大。本研究针对基因型分型测序(genotyping-by-sequencing, GBS)实验流程中的基因组采样偏差,以及各类种群水平数据过滤策略的影响展开评估。本研究选取杨属(Populus)两个物种的数据作为研究对象:该属基因组规模相对较小,正逐渐成为种群基因组学研究的热门靶标类群。本研究通过对毛果杨(Populus trichocarpa)基因组(Nisqually-1)的分析,估算了基因组采样的比例与模式,并证实:在该物种中使用甲基化敏感型限制性内切酶ApeKI时,会出现显著的编码区域富集偏差。本研究还利用近缘物种美洲山杨(P. tremuloides)的种群水平数据,探究了多种GBS数据过滤方案,以保留可用于种群遗传分析的高深度、高信息性单核苷酸多态性(single nucleotide polymorphism, SNPs)。本研究发现,纳入歧义等位基因标注的数据过滤方案,所得到的种群结构与哈迪-温伯格平衡(Hardy-Weinberg equilibrium)指标,与此前基于其他遗传标记对同一种群开展的研究结果最为一致。对过滤后数据(27910个SNPs)的分析结果,同样显示杂合性模式与种群结构与此前利用微卫星标记开展的相关研究高度相似。本研究的应用案例表明,针对自然种群的种群基因组学研究,可便捷地开发出技术与分析层面均较为简便的研究方案。



