Significance Analysis of Prognostic Signatures
收藏资源简介:
A major goal in translational cancer research is to identify biological signatures driving cancer progression and metastasis. A common technique applied in genomics research is to cluster patients using gene expression data from a candidate prognostic gene set, and if the resulting clusters show statistically significant outcome stratification, to associate the gene set with prognosis, suggesting its biological and clinical importance. Recent work has questioned the validity of this approach by showing in several breast cancer data sets that “random” gene sets tend to cluster patients into prognostically variable subgroups. This work suggests that new rigorous statistical methods are needed to identify biologically informative prognostic gene sets. To address this problem, we developed Significance Analysis of Prognostic Signatures (SAPS) which integrates standard prognostic tests with a new prognostic significance test based on stratifying patients into prognostic subtypes with random gene sets. SAPS ensures that a significant gene set is not only able to stratify patients into prognostically variable groups, but is also enriched for genes showing strong univariate associations with patient prognosis, and performs significantly better than random gene sets. We use SAPS to perform a large meta-analysis (the largest completed to date) of prognostic pathways in breast and ovarian cancer and their molecular subtypes. Our analyses show that only a small subset of the gene sets found statistically significant using standard measures achieve significance by SAPS. We identify new prognostic signatures in breast and ovarian cancer and their corresponding molecular subtypes, and we show that prognostic signatures in ER negative breast cancer are more similar to prognostic signatures in ovarian cancer than to prognostic signatures in ER positive breast cancer. SAPS is a powerful new method for deriving robust prognostic biological signatures from clinically annotated genomic datasets.
转化肿瘤学研究(translational cancer research)的核心目标之一,是识别驱动癌症进展与转移的生物学特征(biological signatures)。基因组学研究中常用的一类方法为:基于候选预后基因集的基因表达数据对患者进行聚类,若所得聚类结果呈现出具有统计学显著性的预后分层效果,则将该基因集与预后相关联,以此暗示其生物学与临床重要性。近期有研究对该方法的有效性提出了质疑——其在多项乳腺癌数据集(breast cancer data sets)中证实,“随机”生成的基因集往往也能将患者聚类为预后差异显著的亚组。该研究提示,亟需开发严谨的新型统计方法,以识别具备生物学信息价值的预后基因集。为解决这一问题,我们研发了预后特征显著性分析(Significance Analysis of Prognostic Signatures, SAPS)工具,该方法将标准预后检验与一项基于随机基因集构建预后亚型的新型预后显著性检验相结合。SAPS可确保:具有统计学显著性的基因集不仅能够将患者分层为预后差异显著的组别,还会富集于那些与患者预后存在强单变量关联的基因,且其表现显著优于随机基因集。我们运用SAPS对乳腺癌与卵巢癌及其分子亚型的预后通路开展了迄今为止规模最大的大型荟萃分析(meta-analysis)。分析结果显示:仅少数通过标准检验方法获得统计学显著性的基因集,可经由SAPS验证为显著。我们在乳腺癌与卵巢癌及其对应分子亚型中识别出了全新的预后特征,同时证实:雌激素受体(ER)阴性乳腺癌的预后特征,与卵巢癌的预后特征相似度,要高于其与雌激素受体(ER)阳性乳腺癌预后特征的相似度。SAPS是一款从临床注释基因组数据集(clinically annotated genomic datasets)中获取稳健预后生物学特征的高效全新方法。



