Gene-Centric Characteristics of Genome-Wide Association Studies
收藏资源简介:
BackgroundThe high-throughput genotyping chips have contributed greatly to genome-wide association (GWA) studies to identify novel disease susceptibility single nucleotide polymorphisms (SNPs). The high-density chips are designed using two different SNP selection approaches, the direct gene-centric approach, and the indirect quasi-random SNPs or linkage disequilibrium (LD)-based tagSNPs approaches. Although all these approaches can provide high genome coverage and ascertain variants in genes, it is not clear to which extent these approaches could capture the common genic variants. It is also important to characterize and compare the differences between these approaches.Methodology/Principal FindingsIn our study, by using both the Phase II HapMap data and the disease variants extracted from OMIM, a gene-centric evaluation was first performed to evaluate the ability of the approaches in capturing the disease variants in Caucasian population. Then the distribution patterns of SNPs were also characterized in genic regions, evolutionarily conserved introns and nongenic regions, ontologies and pathways. The results show that, no mater which SNP selection approach is used, the current high-density SNP chips provide very high coverage in genic regions and can capture most of known common disease variants under HapMap frame. The results also show that the differences between the direct and the indirect approaches are relatively small. Both have similar SNP distribution patterns in these gene-centric characteristics.Conclusions/SignificanceThis study suggests that the indirect approaches not only have the advantage of high coverage but also are useful for studies focusing on various functional SNPs either in genes or in the conserved regions that the direct approach supports. The study and the annotation of characteristics will be helpful for designing and analyzing GWA studies that aim to identify genetic risk factors involved in common diseases, especially variants in genes and conserved regions.
研究背景 高通量基因分型芯片极大地推动了全基因组关联研究(Genome-Wide Association Study, GWA)在识别新型疾病易感单核苷酸多态性(Single Nucleotide Polymorphism, SNP)方面的应用。高密度SNP芯片基于两种不同的SNP筛选策略设计:直接以基因为中心的策略,以及间接的准随机SNP策略或基于连锁不平衡(Linkage Disequilibrium, LD)的标签SNP(tagSNP)策略。尽管上述策略均可实现较高的基因组覆盖度并检测基因内的变异,但目前尚不清楚这些策略能够在多大程度上捕获常见的基因内变异,且对不同策略间的差异进行表征与比较也具有重要意义。 研究方法与主要结果 本研究结合使用HapMap第二阶段数据与从在线人类孟德尔遗传(Online Mendelian Inheritance in Man, OMIM)数据库中提取的疾病变异,首先开展以基因为中心的评估,分析上述策略在高加索人群中捕获疾病变异的能力。此外,本研究还对SNP在基因区域、进化保守内含子区、非基因区域以及基因本体(Gene Ontology, GO)和通路中的分布模式进行了表征。研究结果显示,无论采用何种SNP筛选策略,当前的高密度SNP芯片在基因区域均具备极高的覆盖度,且在HapMap框架下可捕获绝大多数已知的常见疾病变异。同时结果还表明,直接策略与间接策略之间的差异相对较小,二者在上述以基因为中心的特征中呈现相似的SNP分布模式。 研究结论与意义 本研究表明,间接策略不仅具备高覆盖度的优势,同时也适用于聚焦于基因内或直接策略所覆盖的保守区域内各类功能性SNP的研究。本研究及其特征注释结果,将有助于设计与分析旨在识别常见疾病相关遗传风险因子(尤其是基因内及保守区域内的变异)的全基因组关联研究。



