MOESM9 of An integrative methodology based on protein-protein interaction networks for identification and functional annotation of disease-relevant genes applied to channelopathies
收藏资源简介:
Additional file 9 Quantitative validation by significance analysis of DAVID search against other phenotype-oriented resources. We searched the nine relevant genes resulted from the workflow in PheGenI [56], ToppGene [57] and g:Profiler [58]. We quantitatively evaluated this search selecting those terms with a significance less than 0.05 using Benjamini-Hochberg FDR statistic. We obtained a minor result in DAVID search (OMIM search did not offer the phenotypes p-values, unlike GAP DISEASE database). Even so, results are useful to develop a quantitative comparison between semiautomatic platforms and bibliographic search systems (sheet 1). From these results we represented the genotype-phenotype association networks to compare easily each p-value phenotype obtained (sheet 2). It should be noted that p-values of clinical phenotypes could be only obtained from one of the two databases explored through DAVID (GAP DISEASE database), and so the genotype-phenotype association network is sparser than the network of the manuscript (section A in Figs. 5, 6, 7). Yet, it is demonstrated that the workflow results are statistically significant and are as valid as or even better than systematic or exhaustive reviews. Then, we created three Boolean tables (in sheets 3, 4, 5) comparing each phenotype obtained from each search; these tables were then converted to binary matrices and clustering multivariate statistical analyses and bootstrap validations were carried out. This approach demonstrated that the results provided in the manuscript, obtained from DAVID (DAVID_m) and systematic and exhaustive reviews, clustered together in a robust and significant way (sheets 3, 4, 5). Hence, this workflow builds as productive results as a non-automatic research but in a quicker way allowing the extraction of information which a priori might not seem relevant when the starting point is a very large group of genes in disease. Moreover, the results obtained using just significant FDR corrected p-values also cluster in particular branches.
补充材料9:通过显著性分析对DAVID数据库(DAVID)检索结果与其他表型导向资源进行定量验证。我们针对工作流程得到的9个相关基因,在PheGenI数据库(PheGenI)[56]、ToppGene数据库(ToppGene)[57]与g:Profiler工具(g:Profiler)[58]中开展了检索。我们采用本雅明尼-霍赫贝格假发现率(Benjamini-Hochberg FDR)统计量,筛选显著性P值小于0.05的表型条目,对本次检索结果进行定量评估。我们在DAVID数据库检索中得到了较为有限的结果(与GAP DISEASE数据库不同,在线人类孟德尔遗传数据库(OMIM)无法提供表型相关P值)。即便如此,本次结果仍可用于半自动平台与文献检索系统之间的定量对比(见工作表1)。基于上述结果,我们构建了基因型-表型关联网络(genotype-phenotype association network),以直观对比所得到的各表型P值(见工作表2)。需注意的是,临床表型的P值仅可通过DAVID检索的两个数据库中的GAP DISEASE数据库获取,因此本次构建的基因型-表型关联网络相较于本文正文的关联网络更为稀疏(对应图5、6、7中的A部分)。但研究结果证实,本工作流程得到的结果具有统计学显著性,其有效性不亚于甚至优于系统性综述或全面综述。随后,我们构建了三张布尔表(Boolean table)(见工作表3、4、5),用于对比各检索得到的表型;随后将这些布尔表转换为二元矩阵(binary matrix),并开展了多元聚类统计分析与Bootstrap验证(Bootstrap validation)。该分析方法证实,本文中来自DAVID(DAVID_m)以及系统性、全面性综述的结果,能够以稳健且显著的方式聚为一类(见工作表3、4、5)。因此,本工作流程能够产出与非自动化研究相当的有效结果,且效率更高,可挖掘出当以疾病领域的大规模基因集作为研究起点时,先验看来似乎无关的信息。此外,仅采用经FDR校正的显著性P值得到的结果,也能在聚类分析中形成特定的分支。



