遇见数据集

A Probabilistic Model to Predict Clinical Phenotypic Traits from Genome Sequencing

收藏
Figshare2016-01-15 更新2026-04-29 收录
官方服务:

资源简介:

Genetic screening is becoming possible on an unprecedented scale. However, its utility remains controversial. Although most variant genotypes cannot be easily interpreted, many individuals nevertheless attempt to interpret their genetic information. Initiatives such as the Personal Genome Project (PGP) and Illumina's Understand Your Genome are sequencing thousands of adults, collecting phenotypic information and developing computational pipelines to identify the most important variant genotypes harbored by each individual. These pipelines consider database and allele frequency annotations and bioinformatics classifications. We propose that the next step will be to integrate these different sources of information to estimate the probability that a given individual has specific phenotypes of clinical interest. To this end, we have designed a Bayesian probabilistic model to predict the probability of dichotomous phenotypes. When applied to a cohort from PGP, predictions of Gilbert syndrome, Graves' disease, non-Hodgkin lymphoma, and various blood groups were accurate, as individuals manifesting the phenotype in question exhibited the highest, or among the highest, predicted probabilities. Thirty-eight PGP phenotypes (26%) were predicted with area-under-the-ROC curve (AUC)>0.7, and 23 (15.8%) of these were statistically significant, based on permutation tests. Moreover, in a Critical Assessment of Genome Interpretation (CAGI) blinded prediction experiment, the models were used to match 77 PGP genomes to phenotypic profiles, generating the most accurate prediction of 16 submissions, according to an independent assessor. Although the models are currently insufficiently accurate for diagnostic utility, we expect their performance to improve with growth of publicly available genomics data and model refinement by domain experts.

基因筛查正以前所未有的规模得以实现,但其应用价值仍存在争议。尽管绝大多数变异基因型尚难以解读,但仍有许多人尝试自行解析其遗传信息。诸如个人基因组计划(Personal Genome Project, PGP)以及因美纳(Illumina)的“了解你的基因组”(Understand Your Genome)等项目,目前已对数千名成年人开展基因组测序,收集表型信息,并开发计算流程以识别每名个体携带的关键变异基因型。此类计算流程会整合数据库、等位基因频率注释信息以及生物信息学分类结果。我们认为,下一步研究方向应为整合各类信息来源,以评估特定个体出现临床关注的特定表型的概率。为此,我们设计了一款贝叶斯概率模型(Bayesian probabilistic model),用于预测二分类表型的发生概率。将该模型应用于PGP队列数据后,吉尔伯特综合征(Gilbert syndrome)、格雷夫斯病(Graves' disease)、非霍奇金淋巴瘤(non-Hodgkin lymphoma)以及多种血型的预测结果均较为准确:携带对应表型的个体,其预测概率均处于最高或接近最高水平。经置换检验验证,在PGP队列的表型预测中,有38项(占比26%)的ROC曲线下面积(area-under-the-ROC curve, AUC)大于0.7,其中23项(占比15.8%)具有统计学显著性。此外,在一场“基因组解读关键评估”(Critical Assessment of Genome Interpretation, CAGI)的盲法预测实验中,研究团队使用该模型将77份PGP基因组与表型特征进行匹配,经独立评估员判定,其预测结果在16份提交方案中准确度最高。尽管当前模型的准确度尚不足以应用于临床诊断,但我们预计,随着公开基因组数据的积累以及领域专家对模型的优化,其性能将得到提升。

创建时间:
2016-01-15
二维码
社区交流群
二维码
科研交流群
商业服务