遇见数据集

GAIN-BRCA: A graph-based explainable AI framework for breast cancer subtype classification based on multi-omics

收藏
Zenodo2025-07-09 更新2026-05-26 收录
官方服务:

资源简介:

Motivation Contextual integration of multiomic datasets from the same patient could improve the accuracy of subtype prediction algorithms to help with better prognosis and management of breast cancer. Previous machine learning models have underexplored the graph-based integration, hence unable to leverage the biological associations among different omics modalities. Here, we developed a graph-based method, GAIN-BRCA, using the native features from mRNA, DNA methylation (CpG), and miRNA data as well as the synthesized features from their interactions. GAIN-BRCA computes weightage from miRNA-mRNA and CpG-mRNA interactions to derive a new transformed feature vector that captures the essential biological context. Results GAIN-BRCA demonstrates superior performance with an AUROC of 0.98. GAIN-BRCA, with an accuracy of 0.92 also outperformed the existing methods like MOGONET and moBRCA-net with accuracies of 0.72 and 0.86, respectively. Kaplan-Meier survival analysis revealed subtype-specific prognostic genes, including KRAS in Luminal A (P value = 0.041), TOX in Luminal B (P value = 0.008), and MITF and TOB1 in HER2+ (P values = 0.029 and 0.025, respectively). However, no single gene demonstrated a significant survival correlation unique to the Basal subtype. GAIN-BRCA framework, in combination with SHAP, has identified several subtype-specific biomarkers to aid in the development of precision therapeutics for breast cancer subtypes. Availability and implementation GAIN-BRCA code is publicly accessible on https://github.com/GudaLab/GAIN-BRCA.

研究动机 对同一患者的多组学数据集进行上下文整合,可提升亚型预测算法的准确性,助力乳腺癌的精准预后与临床管理。既往机器学习模型对基于图的整合方式探索不足,未能充分利用不同组学模态间的生物学关联。本研究开发了一款基于图的方法GAIN-BRCA,该方法采用mRNA、DNA甲基化(CpG)及miRNA数据的原生特征,同时结合它们之间相互作用生成的合成特征。GAIN-BRCA通过计算miRNA-mRNA与CpG-mRNA相互作用的权重,得到可捕捉核心生物学上下文的新型转换特征向量。 实验结果 GAIN-BRCA展现出优异性能,其受试者工作特征曲线下面积(Area Under the Receiver Operating Characteristic Curve,AUROC)达0.98。该模型的分类准确率为0.92,同样优于现有同类方法MOGONET与moBRCA-net(二者准确率分别为0.72与0.86)。卡普兰-迈耶(Kaplan-Meier)生存分析揭示了亚型特异性预后基因,包括Luminal A型中的KRAS(P值=0.041)、Luminal B型中的TOX(P值=0.008),以及HER2阳性型中的MITF与TOB1(P值分别为0.029与0.025)。不过,基底样亚型未发现具有显著生存相关性的独有基因。GAIN-BRCA框架结合SHapley加性解释(SHAP),已识别出多种亚型特异性生物标志物,可为乳腺癌各亚型的精准治疗药物研发提供辅助支撑。 可用性与实现 GAIN-BRCA的代码已公开上传至https://github.com/GudaLab/GAIN-BRCA。

提供机构:
Zenodo
创建时间:
2025-04-08
二维码
社区交流群
二维码
科研交流群
商业服务