A Model-Based Joint Identification of Differentially Expressed Genes and Phenotype-Associated Genes
收藏资源简介:
Over the last decade, many analytical methods and tools have been developed for microarray data. The detection of differentially expressed genes (DEGs) among different treatment groups is often a primary purpose of microarray data analysis. In addition, association studies investigating the relationship between genes and a phenotype of interest such as survival time are also popular in microarray data analysis. Phenotype association analysis provides a list of phenotype-associated genes (PAGs). However, it is sometimes necessary to identify genes that are both DEGs and PAGs. We consider the joint identification of DEGs and PAGs in microarray data analyses. The first approach we used was a naïve approach that detects DEGs and PAGs separately and then identifies the genes in an intersection of the list of PAGs and DEGs. The second approach we considered was a hierarchical approach that detects DEGs first and then chooses PAGs from among the DEGs or vice versa. In this study, we propose a new model-based approach for the joint identification of DEGs and PAGs. Unlike the previous two-step approaches, the proposed method identifies genes simultaneously that are DEGs and PAGs. This method uses standard regression models but adopts different null hypothesis from ordinary regression models, which allows us to perform joint identification in one-step. The proposed model-based methods were evaluated using experimental data and simulation studies. The proposed methods were used to analyze a microarray experiment in which the main interest lies in detecting genes that are both DEGs and PAGs, where DEGs are identified between two diet groups and PAGs are associated with four phenotypes reflecting the expression of leptin, adiponectin, insulin-like growth factor 1, and insulin. Model-based approaches provided a larger number of genes, which are both DEGs and PAGs, than other methods. Simulation studies showed that they have more power than other methods. Through analysis of data from experimental microarrays and simulation studies, the proposed model-based approach was shown to provide a more powerful result than the naïve approach and the hierarchical approach. Since our approach is model-based, it is very flexible and can easily handle different types of covariates.
近十年来,针对微阵列数据(microarray data)已涌现出诸多分析方法与工具。不同处理组间差异表达基因(differentially expressed genes, DEGs)的检测,往往是微阵列数据分析的核心目标之一。此外,探究基因与目标表型(如生存时间)之间关联的关联研究,在微阵列数据分析中也颇为流行。表型关联分析可得到表型关联基因(phenotype-associated genes, PAGs)列表。但实际研究中有时需要同时识别既是差异表达基因又是表型关联基因的基因。本研究聚焦微阵列数据分析中差异表达基因与表型关联基因的联合识别问题。 我们首先采用了两种经典方法:其一为朴素方法,即分别检测差异表达基因与表型关联基因,再取二者列表的交集以筛选目标基因;其二为分层方法,即先检测差异表达基因,再从差异表达基因中筛选表型关联基因,反之亦可。 本研究提出一种全新的基于模型的差异表达基因与表型关联基因联合识别方法。与前述两步法不同,所提方法可同时识别兼具差异表达基因与表型关联基因属性的基因。该方法基于标准回归模型,但采用了与普通回归模型截然不同的原假设,从而可在单一步骤中完成联合识别。 我们通过实验数据与模拟实验对所提基于模型的方法进行了性能评估。将该方法应用于一项微阵列实验分析,该实验的核心目标是检测同时属于两组饮食处理差异表达基因、且与四种表型相关的基因——这四种表型分别为瘦素(leptin)、脂联素(adiponectin)、胰岛素样生长因子1(insulin-like growth factor 1)以及胰岛素的表达水平。 相较于其他方法,基于模型的方法可筛选出更多同时兼具差异表达基因与表型关联基因属性的基因。模拟实验结果表明,该方法拥有更高的统计效力。通过对实验微阵列数据与模拟数据的分析,我们证实所提基于模型的方法相较于朴素方法与分层方法,可获得更优异的检测效力。由于本方法基于模型构建,因此具备极强的灵活性,可轻松适配不同类型的协变量(covariates)。



