Predictive Power Estimation Algorithm (PPEA) - A New Algorithm to Reduce Overfitting for Genomic Biomarker Discovery
收藏资源简介:
Toxicogenomics promises to aid in predicting adverse effects, understanding the mechanisms of drug action or toxicity, and uncovering unexpected or secondary pharmacology. However, modeling adverse effects using high dimensional and high noise genomic data is prone to over-fitting. Models constructed from such data sets often consist of a large number of genes with no obvious functional relevance to the biological effect the model intends to predict that can make it challenging to interpret the modeling results. To address these issues, we developed a novel algorithm, Predictive Power Estimation Algorithm (PPEA), which estimates the predictive power of each individual transcript through an iterative two-way bootstrapping procedure. By repeatedly enforcing that the sample number is larger than the transcript number, in each iteration of modeling and testing, PPEA reduces the potential risk of overfitting. We show with three different cases studies that: (1) PPEA can quickly derive a reliable rank order of predictive power of individual transcripts in a relatively small number of iterations, (2) the top ranked transcripts tend to be functionally related to the phenotype they are intended to predict, (3) using only the most predictive top ranked transcripts greatly facilitates development of multiplex assay such as qRT-PCR as a biomarker, and (4) more importantly, we were able to demonstrate that a small number of genes identified from the top-ranked transcripts are highly predictive of phenotype as their expression changes distinguished adverse from nonadverse effects of compounds in completely independent tests. Thus, we believe that the PPEA model effectively addresses the over-fitting problem and can be used to facilitate genomic biomarker discovery for predictive toxicology and drug responses.
毒理基因组学(Toxicogenomics)有望辅助预测不良反应、阐明药物作用或毒性机制,以及发掘意外或继发性药理学效应。然而,利用高维度、高噪声基因组数据构建不良反应预测模型时,极易出现过拟合现象。基于此类数据集构建的模型,往往包含大量与模型拟预测的生物学效应无明确功能关联的基因,这会给模型结果的解读带来极大挑战。为解决上述问题,我们开发了一种全新算法——预测效能估计算法(Predictive Power Estimation Algorithm,PPEA),该算法通过迭代双向自助抽样流程评估单个转录本的预测效能。通过在每一轮建模与测试迭代中始终保证样本量大于转录本(transcript)数量,PPEA可降低过拟合的潜在风险。我们通过三项不同的案例研究证实:(1)PPEA可在相对较少的迭代次数内,快速得到单个转录本预测效能的可靠排序;(2)排名靠前的转录本往往与其拟预测的表型存在功能关联;(3)仅使用预测效能最优的排名靠前转录本,可极大推动以定量实时聚合酶链反应(qRT-PCR)为代表的多重检测技术作为生物标志物的开发;(4)更为关键的是,我们证实从排名靠前的转录本中筛选出的少量基因,具备极强的表型预测能力——在完全独立的测试中,这些基因的表达变化可有效区分化合物的不良反应与非不良反应。因此,我们认为PPEA模型可有效解决过拟合问题,能够助力预测毒理学与药物响应相关的基因组学生物标志物发掘工作。



