遇见数据集

Identifying In-<em>Trans</em> Process Associated Genes in Breast Cancer by Integrated Analysis of Copy Number and Expression Data

收藏
NIAID Data Ecosystem2026-03-07 收录
官方服务:

资源简介:

Genomic copy number alterations are common in cancer. Finding the genes causally implicated in oncogenesis is challenging because the gain or loss of a chromosomal region may affect a few key driver genes and many passengers. Integrative analyses have opened new vistas for addressing this issue. One approach is to identify genes with frequent copy number alterations and corresponding changes in expression. Several methods also analyse effects of transcriptional changes on known pathways. Here, we propose a method that analyses in-cis correlated genes for evidence of in-trans association to biological processes, with no bias towards processes of a particular type or function. The method aims to identify cis-regulated genes for which the expression correlation to other genes provides further evidence of a network-perturbing role in cancer. The proposed unsupervised approach involves a sequence of statistical tests to systematically narrow down the list of relevant genes, based on integrative analysis of copy number and gene expression data. A novel adjustment method handles confounding effects of co-occurring copy number aberrations, potentially a large source of false positives in such studies. Applying the method to whole-genome copy number and expression data from 100 primary breast carcinomas, 6373 genes were identified as commonly aberrant, 578 were highly in-cis correlated, and 56 were in addition associated in-trans to biological processes. Among these in-trans process associated and cis-correlated (iPAC) genes, 28% have previously been reported as breast cancer associated, and 64% as cancer associated. By combining statistical evidence from three separate subanalyses that focus respectively on copy number, gene expression and the combination of the two, the proposed method identifies several known and novel cancer driver candidates. Validation in an independent data set supports the conclusion that the method identifies genes implicated in cancer.

基因组拷贝数变异(genomic copy number alterations)在癌症中极为常见。寻找在肿瘤发生(oncogenesis)中具有因果关联的基因颇具挑战,这是因为染色体区域的扩增或缺失可能仅影响少数关键驱动基因(driver genes),以及大量过客基因(passenger genes)。整合分析为解决该问题开辟了全新研究视角。一类研究思路是识别发生频繁拷贝数变异且伴随表达水平相应改变的基因。另有部分方法会分析转录组变化对已知通路的影响。本文提出一种分析顺式(cis)相关基因的方法,用于寻找其与生物过程存在反式(trans)关联的证据,且不会对特定类型或功能的通路产生偏倚。该方法旨在筛选顺式调控基因,这类基因与其他基因的表达相关性,可为其在癌症中发挥网络扰动作用提供进一步佐证。所提出的无监督方法包含一系列统计检验步骤,基于对拷贝数与基因表达数据的整合分析,系统性地缩小相关基因的候选列表。一种全新的校正方法可处理共现拷贝数异常带来的混杂效应——这类混杂效应是此类研究中假阳性结果的重要潜在来源。我们将该方法应用于100例原发性乳腺癌的全基因组拷贝数与基因表达数据,共鉴定出6373个常见异常基因,其中578个具有高度顺式相关性,另有56个同时与生物过程存在反式关联。在这些顺式相关且与生物过程存在反式关联(in-trans process associated and cis-correlated, 简称iPAC)基因中,28%此前已被报道与乳腺癌相关,64%与癌症相关。通过整合三项分别聚焦拷贝数、基因表达及二者联合的子分析的统计证据,所提方法鉴定出多个已知与新型癌症驱动候选基因。在独立数据集上的验证结果证实,该方法可筛选出与癌症相关的基因。

创建时间:
2016-01-18
二维码
社区交流群
二维码
科研交流群
商业服务