遇见数据集

Multi-study Integration of Brain Cancer Transcriptomes Reveals Organ-Level Molecular Signatures

收藏
Figshare2016-01-18 更新2026-04-29 收录
官方服务:

资源简介:

We utilized abundant transcriptomic data for the primary classes of brain cancers to study the feasibility of separating all of these diseases simultaneously based on molecular data alone. These signatures were based on a new method reported herein – Identification of Structured Signatures and Classifiers (ISSAC) – that resulted in a brain cancer marker panel of 44 unique genes. Many of these genes have established relevance to the brain cancers examined herein, with others having known roles in cancer biology. Analyses on large-scale data from multiple sources must deal with significant challenges associated with heterogeneity between different published studies, for it was observed that the variation among individual studies often had a larger effect on the transcriptome than did phenotype differences, as is typical. For this reason, we restricted ourselves to studying only cases where we had at least two independent studies performed for each phenotype, and also reprocessed all the raw data from the studies using a unified pre-processing pipeline. We found that learning signatures across multiple datasets greatly enhanced reproducibility and accuracy in predictive performance on truly independent validation sets, even when keeping the size of the training set the same. This was most likely due to the meta-signature encompassing more of the heterogeneity across different sources and conditions, while amplifying signal from the repeated global characteristics of the phenotype. When molecular signatures of brain cancers were constructed from all currently available microarray data, 90% phenotype prediction accuracy, or the accuracy of identifying a particular brain cancer from the background of all phenotypes, was found. Looking forward, we discuss our approach in the context of the eventual development of organ-specific molecular signatures from peripheral fluids such as the blood.

本研究针对主要类型的脑肿瘤,利用海量转录组数据,探究仅依托分子数据即可同时区分所有此类肿瘤的可行性。本研究提出的结构化特征与分类器识别法(Identification of Structured Signatures and Classifiers,ISSAC),可生成包含44个独特基因的脑肿瘤标志物组合。其中多数基因已被证实与本研究涉及的脑肿瘤存在明确关联,其余基因则在癌症生物学中具有已知功能。 针对多来源大规模数据开展分析时,需应对不同已发表研究间异质性带来的显著挑战:正如常见情形,单个研究间的差异对转录组的影响往往大于表型差异带来的影响。据此,本研究仅纳入每种表型至少包含两项独立研究的案例,并采用统一的预处理流程对所有研究的原始数据进行重处理。 本研究发现,即便保持训练集规模不变,基于多数据集学习得到的特征可显著提升独立验证集上预测性能的可重复性与准确性。这一现象大概率源于元特征(meta-signature)覆盖了不同来源与实验条件下的更多异质性,同时放大了表型重复出现的全局特征所携带的信号。当利用当前所有可用的微阵列(microarray)数据构建脑肿瘤分子特征时,实现了90%的表型预测准确率——即从所有表型背景中识别特定脑肿瘤的准确率。 展望未来,我们将在通过外周体液(如血液)最终构建器官特异性分子特征的研究脉络中,探讨本研究提出的方法。

创建时间:
2016-01-18
二维码
社区交流群
二维码
科研交流群
商业服务