FunSPU: A versatile and adaptive multiple functional annotation-based association test of whole-genome sequencing data
收藏资源简介:
Despite ongoing large-scale population-based whole-genome sequencing (WGS) projects such as the NIH NHLBI TOPMed program and the NHGRI Genome Sequencing Program, WGS-based association analysis of complex traits remains a tremendous challenge due to the large number of rare variants, many of which are non-trait-associated neutral variants. External biological knowledge, such as functional annotations based on the ENCODE, Epigenomics Roadmap and GTEx projects, may be helpful in distinguishing causal rare variants from neutral ones; however, each functional annotation can only provide certain aspects of the biological functions. Our knowledge for selecting informative annotations a priori is limited, and incorporating non-informative annotations will introduce noise and lose power. We propose FunSPU, a versatile and adaptive test that incorporates multiple biological annotations and is adaptive at both the annotation and variant levels and thus maintains high power even in the presence of noninformative annotations. In addition to extensive simulations, we illustrate our proposed test using the TWINSUK cohort (n = 1,752) of UK10K WGS data based on six functional annotations: CADD, RegulomeDB, FunSeq, Funseq2, GERP++, and GenoSkyline. We identified genome-wide significant genetic loci on chromosome 19 near gene TOMM40 and APOC4-APOC2 associated with low-density lipoprotein (LDL), which are replicated in the UK10K ALSPAC cohort (n = 1,497). These replicated LDL-associated loci were missed by existing rare variant association tests that either ignore external biological information or rely on a single source of biological knowledge. We have implemented the proposed test in an R package “FunSPU”.
尽管诸如美国国立卫生研究院(National Institutes of Health, NIH)下属国家心肺血液研究所(National Heart, Lung, and Blood Institute, NHLBI)的TOPMed计划以及美国国家人类基因组研究所(National Human Genome Research Institute, NHGRI)基因组测序计划等大规模人群全基因组测序(Whole-Genome Sequencing, WGS)项目正在持续推进,但基于WGS的复杂性状关联分析仍面临巨大挑战:罕见变异数量庞大,且其中多数为与性状无关的中性变异。外部生物学知识,例如基于ENCODE、表观基因组路线图(Epigenomics Roadmap)和GTEx项目(Genotype-Tissue Expression Project)的功能注释,或可助力区分致病变异与中性变异;但每一类功能注释仅能反映生物学功能的某一侧面。我们在先验筛选有信息量的注释时存在认知局限,若引入无信息的注释,反而会引入噪声并降低检验效能。为此我们提出FunSPU方法,这是一种通用且自适应的检验工具,可整合多种生物学注释,且能在注释与变异两个层面实现自适应,因此即便存在无信息注释,仍可保持较高的检验效力。除开展了大量模拟实验外,我们还基于6种功能注释——包括CADD、RegulomeDB、FunSeq、Funseq2、GERP++以及GenoSkyline,利用UK10K全基因组测序数据中的TWINSUK队列(n=1752)对所提方法进行了实例演示。最终在19号染色体TOMM40基因附近及APOC4-APOC2区域,鉴定到与低密度脂蛋白(Low-Density Lipoprotein, LDL)相关的全基因组显著性遗传位点,并在UK10K的ALSPAC队列(n=1497)中完成了重复验证。现有罕见变异关联检验方法要么忽略外部生物学信息,要么仅依赖单一生物学知识来源,因此未能发现这些与LDL相关的可重复位点。我们已将所提出的检验方法实现于R包"FunSPU"中。



