遇见数据集

A scalable, accurate, and universal analysis framework using individual-level allele frequency for large-scale genetic association studies in an admixed population

收藏
Zenodo2023-09-17 更新2026-05-26 收录
数据链接:
官方服务:

资源简介:

Inclusion of individuals with diverse or admixed genetic ancestries is crucial to discover novel findings that may be missed by genomics analyses rooted solely in Caucasian population. Here, we present an analysis framework, SPAmix, which is scalable to a large-scale biobank data analysis including hundreds of thousands of admixed individuals and is universally applicable to various types of complex traits including binary trait, quantitative trait, time-to-event trait, longitudinal traits, etc. For each genetic variant, SPAmix uses genotype data and genetic principal components (PCs) to estimate individual-level allele frequency, which is subsequently used to calibrate p values via a retrospective analysis. A hybrid strategy including saddlepoint approximation (SPA) can greatly increase the accuracy to analyze rare genetic variants, especially if the phenotypic distribution is unbalanced or extremely unbalanced. Compared to Tractor, SPAmix does not require local ancestry information and can be straightforwardly applicable to a multi-way admixed population. Meanwhile, SPAmix can also be extended to SPAmix<sub>local</sub> in which the local ancestry can be incorporated if available. In addition, we propose SPAmix<sub>CCT</sub> to combine the p values of SPAmix and SPAmix<sub>local</sub> via Cauchy combination (CCT). SPAmix<sub>local</sub> performs close to Tractor when analyzing quantitative traits and is more accurate when analyzing binary traits with an unbalanced case-control ratio. And SPAmix<sub>CCT </sub>is an optimal unified approach for various cross-ancestry genetic architectures. Extensive simulation studies and real data analyses of 369,314 UK Biobank individuals from multiple ancestries demonstrated that SPAmix is scalable and can discover novel hits while controlling type I error rates well.

纳入具有多样化或混合遗传祖先的个体,对于发现仅基于高加索人群的基因组学分析可能遗漏的全新研究发现至关重要。在此,我们提出一款分析框架SPAmix,其可扩展至包含数十万混合遗传祖先个体的大规模生物样本库数据分析,且可普遍适用于多种复杂性状类型,包括二分类性状、数量性状、生存时间性状、纵向性状等。针对每个遗传变异,SPAmix利用基因型数据与遗传主成分(genetic principal components, PCs)估计个体水平的等位基因频率,随后通过回顾性分析对P值进行校准。包含鞍点近似(saddlepoint approximation, SPA)的混合策略可大幅提升稀有遗传变异的分析准确性,尤其在表型分布失衡或极度失衡的场景下。与Tractor相比,SPAmix无需本地祖先信息,可直接应用于多向混合遗传祖先人群。与此同时,SPAmix还可扩展为SPAmix<sub>local</sub>,在可获取本地祖先信息时可将其纳入分析。此外,我们提出SPAmix<sub>CCT</sub>,通过柯西组合(Cauchy combination, CCT)整合SPAmix与SPAmix<sub>local</sub>的P值。SPAmix<sub>local</sub>在分析数量性状时的性能与Tractor相近,而在分析病例对照比失衡的二分类性状时则具备更高的准确性。SPAmix<sub>CCT</sub>则是适配多种跨祖先遗传结构的最优统一分析方法。通过对来自多个祖先群体的369314名英国生物样本库(UK Biobank)个体开展大规模模拟研究与真实数据分析,结果表明SPAmix具备良好的可扩展性,可在严格控制一类错误率的同时发现全新的遗传关联位点。

提供机构:
Zenodo
创建时间:
2023-09-14
二维码
社区交流群
二维码
科研交流群
商业服务