遇见数据集

Data from: Parallel tagged amplicon sequencing reveals major lineages and phylogenetic structure in the North American tiger salamander (Ambystoma tigrinum) species complex

收藏
DataONE2012-08-30 更新2024-06-27 收录
数据链接:
官方服务:

资源简介:

Modern analytical methods for population genetics and phylogenetics are expected to provide more accurate results when data from multiple genome-wide loci are analyzed. We present the results of an initial application of parallel tagged sequencing (PTS) on a next generation platform to sequence thousands of barcoded PCR amplicons generated from 95 nuclear loci and 93 individuals sampled across the range of the tiger salamander (Ambystoma tigrinum) species complex. To manage the bioinformatic processing of this large data set (344,330 reads), we developed a pipeline that sorts PTS data by barcode and locus, identifies high-quality variable nucleotides, and yields phased haplotype sequences for each individual at each locus. Our sequencing and bioinformatic strategy resulted in a genome-wide data set with relatively low levels of missing data and a wide range of nucleotide variation. STRUCTURE analyses of these data in a genotypic format resulted in strongly supported assignments for the majority of individuals into nine geographically defined genetic clusters. Species tree analyses of the most variable loci using a multi-species coalescent model resulted in strong support for most branches in the species tree; however, analyses including more than 50 loci produced parameter sampling trends that indicated a lack of convergence on the posterior distribution. Overall, these results demonstrate the potential for amplicon-based PTS to rapidly generate large-scale data for population genetic and phylogenetic-based research.

当分析全基因组多位点数据时,群体遗传学与系统发育学的现代分析方法有望获得更为精准的研究结果。本研究展示了平行标记测序(parallel tagged sequencing, PTS)在新一代测序平台上的首次应用:该平台对源自95个核基因座、采自虎钝口螈(Ambystoma tigrinum)物种复合群分布范围内的93个个体所扩增得到的数千个带条形码的PCR扩增子进行了测序。为处理该包含344330条读段的大型数据集的生物信息学分析工作,我们开发了一套分析管线,其可按条形码与基因座对PTS数据进行分选,识别高质量的可变核苷酸,并为每个个体的每个基因座生成阶段性单倍型序列。本研究采用的测序与生物信息学策略所构建的全基因组数据集,缺失数据占比较低,且涵盖了广泛范围的核苷酸变异。以基因型格式对这些数据开展的STRUCTURE分析,将绝大多数个体明确归属为9个地理界定的遗传集群,且支持度极高。采用多物种溯祖模型对变异度最高的基因座开展的物种树分析,为物种树的绝大多数分支提供了极强的支持;然而,当纳入超过50个基因座进行分析时,参数采样趋势显示其后验分布未达到收敛。总体而言,本研究结果证实了基于扩增子的平行标记测序技术,能够快速为群体遗传学与系统发育学相关研究生成大规模数据的潜力。

创建时间:
2012-08-30
二维码
社区交流群
二维码
科研交流群
商业服务