Data from: A draft fur seal genome provides insights into factors affecting SNP validation and how to mitigate them
收藏资源简介:
Custom genotyping arrays provide a flexible and accurate means of genotyping single nucleotide polymorphisms (SNPs) in a large number of individuals of essentially any organism. However, validation rates, defined as the proportion of putative SNPs that are verified to be polymorphic in a population, are often very low. A number of potential causes of assay failure have been identified, but none have been explored systematically. In particular, as SNPs are often developed from transcriptomes, parameters relating to the genomic context are rarely taken into account. Here, we assembled a draft Antarctic fur seal (Arctocephalus gazella) genome (assembly size: 2.41Gb; scaffold/contig N50: 3.1Mb/27.5kb). We then used this resource to map the probe sequences of 144 putative SNPs genotyped in 480 individuals. The number of probe-to-genome mappings and alignment length together explained almost a third of the variation in validation success, indicating that sequence uniqueness and proximity to intron-exon boundaries play an important role. The same pattern was found after mapping the probe sequences to the Walrus and Weddell seal genomes, suggesting that the genomes of species divergent by as much as 23 million years can hold information relevant to SNP validation outcomes. Additionally, re-analysis of genotyping data from seven previous studies found the same two variables to be significantly associated with SNP validation success across a variety of taxa. Finally, our study reveals considerable scope for validation rates to be improved, either by simply filtering for SNPs whose flanking sequences align uniquely and completely to a reference genome, or through predictive modeling.
定制基因分型阵列可为几乎所有物种的大量个体开展单核苷酸多态性(single nucleotide polymorphisms, SNPs)分型提供灵活且精准的技术手段。然而,其验证率——即经证实可在目标群体中呈多态性的推定单核苷酸多态性位点所占比例——通常极低。目前已鉴定出多种可能导致分型实验失败的诱因,但尚未有研究对其开展系统性探究。尤为值得关注的是,由于单核苷酸多态性位点往往从转录组中开发获得,与基因组背景相关的参数极少被纳入实验设计考量。本研究首先组装得到南极海狗(Arctocephalus gazella)的草图基因组(组装规模:2.41Gb;支架(scaffold)N50:3.1Mb,重叠群(contig)N50:27.5kb)。随后利用该基因组组装结果,对已在480个个体中完成基因分型的144条推定单核苷酸多态性位点的探针序列进行基因组比对。探针的基因组比对数量与比对长度共同解释了近三分之一的验证成功率变异,表明序列唯一性以及与内含子-外显子边界的邻近性对验证成功率具有重要影响。将探针序列比对至海象和威德尔海豹的基因组后,得到了一致的结果,这表明即使是分化时长长达2300万年的物种,其基因组也可提供与单核苷酸多态性位点验证结果相关的有效信息。此外,对7项既往研究的基因分型数据进行重新分析后发现,上述两个变量在多个类群中均与单核苷酸多态性位点验证成功率呈显著相关。最后,本研究证实可通过两种方式显著提升验证率:一是仅筛选侧翼序列可唯一且完整比对至参考基因组的单核苷酸多态性位点,二是借助预测建模实现精准优化。



