A Unified Method for Detecting Secondary Trait Associations with Rare Variants: Application to Sequence Data
收藏资源简介:
Next-generation sequencing has made possible the detection of rare variant (RV) associations with quantitative traits (QT). Due to high sequencing cost, many studies can only sequence a modest number of selected samples with extreme QT. Therefore association testing in individual studies can be underpowered. Besides the primary trait, many clinically important secondary traits are often measured. It is highly beneficial if multiple studies can be jointly analyzed for detecting associations with commonly measured traits. However, analyzing secondary traits in selected samples can be biased if sample ascertainment is not properly modeled. Some methods exist for analyzing secondary traits in selected samples, where some burden tests can be implemented. However p-values can only be evaluated analytically via asymptotic approximations, which may not be accurate. Additionally, potentially more powerful sequence kernel association tests, variable selection-based methods, and burden tests that require permutations cannot be incorporated. To overcome these limitations, we developed a unified method for analyzing secondary trait associations with RVs (STAR) in selected samples, incorporating all RV tests. Statistical significance can be evaluated either through permutations or analytically. STAR makes it possible to apply more powerful RV tests to analyze secondary trait associations. It also enables jointly analyzing multiple cohorts ascertained under different study designs, which greatly boosts power. The performance of STAR and commonly used RV association tests were comprehensively evaluated using simulation studies. STAR was also implemented to analyze a dataset from the SardiNIA project where samples with extreme low-density lipoprotein levels were sequenced. A significant association between LDLR and systolic blood pressure was identified, which is supported by pharmacogenetic studies. In summary, for sequencing studies, STAR is an important tool for detecting secondary-trait RV associations.
新一代测序技术已使得针对罕见变异(rare variant,RV)与数量性状(quantitative trait,QT)的关联分析成为可能。受限于高昂的测序成本,多数研究仅能对携带极端数量性状的精选样本开展规模有限的测序,因此单一项研究中的关联检验往往功效不足。除核心性状外,研究中通常还会采集多项具有临床重要性的次级性状,若能对多项研究开展联合分析以检测与通用测量性状的关联,将极大提升研究价值。但倘若未对样本招募筛选过程进行恰当建模,针对精选样本的次级性状分析可能会引入偏倚。目前已有部分面向精选样本次级性状的分析方法,可实现部分负担检验(burden test)的应用,但此类方法的P值仅能通过渐近近似的解析方式计算,结果未必精准。此外,部分功效更优的序列核关联检验(sequence kernel association test)、基于变量选择的方法,以及需要置换检验的负担检验均无法被纳入此类分析框架。为克服上述局限,本研究开发了一款针对精选样本中罕见变异次级性状关联分析的统一方法(STAR),可整合所有罕见变异关联检验手段。该方法的统计学显著性可通过置换检验或解析计算两种方式进行评估。STAR可支持使用功效更优的罕见变异关联检验方法开展次级性状关联分析,同时还支持对采用不同研究设计招募的多队列数据开展联合分析,大幅提升检验功效。本研究通过模拟试验全面评估了STAR与主流罕见变异关联检验方法的性能表现。此外,本研究将STAR应用于SardiNIA项目的数据集,该数据集针对极端低密度脂蛋白水平的样本完成了测序,分析结果鉴定出LDLR基因与收缩压之间存在显著关联,该结论得到了药物遗传学研究的佐证。综上,针对测序研究而言,STAR是检测次级性状与罕见变异关联的一款重要工具。



