遇见数据集

Rare Variants Association Analysis in Large-Scale Sequencing Studies at the Single Locus Level

收藏
Figshare2016-09-28 更新2026-04-29 收录
官方服务:

资源简介:

Genetic association analyses of rare variants in next-generation sequencing (NGS) studies are fundamentally challenging due to the presence of a very large number of candidate variants at extremely low minor allele frequencies. Recent developments often focus on pooling multiple variants to provide association analysis at the gene instead of the locus level. Nonetheless, pinpointing individual variants is a critical goal for genomic researches as such information can facilitate the precise delineation of molecular mechanisms and functions of genetic factors on diseases. Due to the extreme rarity of mutations and high-dimensionality, significances of causal variants cannot easily stand out from those of noncausal ones. Consequently, standard false-positive control procedures, such as the Bonferroni and false discovery rate (FDR), are often impractical to apply, as a majority of the causal variants can only be identified along with a few but unknown number of noncausal variants. To provide informative analysis of individual variants in large-scale sequencing studies, we propose the Adaptive False-Negative Control (AFNC) procedure that can include a large proportion of causal variants with high confidence by introducing a novel statistical inquiry to determine those variants that can be confidently dispatched as noncausal. The AFNC provides a general framework that can accommodate for a variety of models and significance tests. The procedure is computationally efficient and can adapt to the underlying proportion of causal variants and quality of significance rankings. Extensive simulation studies across a plethora of scenarios demonstrate that the AFNC is advantageous for identifying individual rare variants, whereas the Bonferroni and FDR are exceedingly over-conservative for rare variants association studies. In the analyses of the CoLaus dataset, AFNC has identified individual variants most responsible for gene-level significances. Moreover, single-variant results using the AFNC have been successfully applied to infer related genes with annotation information.

下一代测序(next-generation sequencing, NGS)研究中的罕见变异遗传关联分析,本质上极具挑战:候选变异数量庞大,且次要等位基因频率极低。现有研究多聚焦于对多个变异进行合并分析,以在基因层面而非位点层面开展关联研究。然而,精准定位单个变异仍是基因组学研究的核心目标——此类信息可助力精准阐明遗传因素对疾病的分子机制与功能。由于突变极度稀有且数据维度极高,致病变异的显著性难以从非致病变异中凸显出来。因此,诸如邦费罗尼校正(Bonferroni)与错误发现率(false discovery rate, FDR)这类标准假阳性控制方法往往难以实际应用,原因在于大多数致病变异仅能与少量且数量未知的非致病变异一同被检出。为实现大规模测序研究中单个变异的有效分析,我们提出了自适应假阴性控制(Adaptive False-Negative Control, AFNC)方法:通过引入全新的统计判定规则,识别可被可靠归类为非致病变异的位点,从而以高置信度覆盖绝大多数致病变异。AFNC 提供了通用分析框架,可适配多种模型与显著性检验方法。该方法计算效率优异,且可自适应调整以适配致病变异的潜在占比与显著性排序质量。针对多种场景开展的大量模拟研究表明,AFNC在识别单个罕见变异方面优势显著,而邦费罗尼校正与FDR在罕见变异关联研究中则过于保守。在对CoLaus数据集的分析中,AFNC成功识别出了对基因层面显著性贡献最大的单个变异。此外,基于AFNC得到的单变异分析结果已被成功用于结合注释信息推断相关基因。

创建时间:
2016-09-28
二维码
社区交流群
二维码
科研交流群
商业服务