Using Network Methodology to Infer Population Substructure
收藏资源简介:
One of the main caveats of association studies is the possible affection by bias due to population stratification. Existing methods rely on model-based approaches like structure and ADMIXTURE or on principal component analysis like EIGENSTRAT. Here we provide a novel visualization technique and describe the problem of population substructure from a graph-theoretical point of view. We group the sequenced individuals into triads, which depict the relational structure, on the basis of a predefined pairwise similarity measure. We then merge the triads into a network and apply community detection algorithms in order to identify homogeneous subgroups or communities, which can further be incorporated as covariates into logistic regression. We apply our method to populations from different continents in the 1000 Genomes Project and evaluate the type 1 error based on the empirical p-values. The application to 1000 Genomes data suggests that the network approach provides a very fine resolution of the underlying ancestral population structure. Besides we show in simulations, that in the presence of discrete population structures, our developed approach maintains the type 1 error more precisely than existing approaches.
关联研究的核心局限之一,在于其可能受到人群分层(population stratification)带来的偏倚干扰。现有方法多采用两类策略:一类是基于模型的方法,如structure与ADMIXTURE;另一类是基于主成分分析的方法,如EIGENSTRAT。本研究提出一种全新的可视化技术,并从图论视角阐释人群亚结构问题。我们基于预设的两两相似性度量,将测序个体划分为三联体(triads)以刻画其关联结构。随后将这些三联体整合为网络,并应用社区检测算法以识别同质亚群或社区;所得结果可进一步作为协变量纳入逻辑回归模型。我们将该方法应用于千人基因组计划(1000 Genomes Project)中来自不同大洲的人群样本,并基于经验p值评估一类错误(type 1 error)率。针对千人基因组计划数据的应用结果表明,该网络方法可精细解析样本潜在的祖先人群结构。此外,模拟实验结果显示,当存在离散人群结构时,本研究所提出的方法相比现有手段,能更精准地控制一类错误率。



