TCGA-SurvReport
收藏资源简介:
TCGA-SurvReport是一个由香港科技大学构建的综合性癌症生存预测基准数据集,整合了六个TCGA队列的病理、临床和分子证据,形成统一的患者报告文本。该数据集包含大量右删失样本(超过60%),旨在为报告为中心的生存分析提供标准化评估平台。数据集的创建通过从TCGA获取原始记录,并结构化组织为可供大语言模型直接处理的报告格式。该数据集用于评估生存预测模型的排序一致性,尤其适用于解决删失数据下的相对预后排序问题,推动语言模型在临床预后中的可靠应用。
TCGA-SurvReport is a comprehensive cancer survival prediction benchmark dataset constructed by The Hong Kong University of Science and Technology, which integrates pathological, clinical and molecular evidence from six TCGA cohorts to form unified patient report texts. This dataset contains a large number of right-censored samples (over 60%), and aims to provide a standardized evaluation platform for report-centric survival analysis. The dataset is created by acquiring original records from TCGA and structuring them into report formats that can be directly processed by large language models. It is used to evaluate the ranking consistency of survival prediction models, and is particularly suitable for solving the relative prognostic ranking problem under censored data, so as to promote the reliable application of language models in clinical prognosis.
CACSurv 数据集概述
CACSurv 是一个与癌症生存预测研究相关的数据集,其全称为 Concordance-Aligned Comparative Learning with Large Language Models for Cancer Survival Prediction,即“基于大型语言模型的一致性对齐比较学习用于癌症生存预测”。该数据集是论文《CACSurv: Concordance-Aligned Comparative Learning with Large Language Models for Cancer Survival Prediction》的官方配套资源。
该数据集的详情页面地址为:https://github.com/xmed-lab/CACSurv
目前,该数据集的代码、模型以及数据集本身尚未公开发布,将在后续版本中逐步放出。因此,当前阶段该数据集暂不可直接获取或使用。
如需了解更多信息,建议持续关注该 GitHub 仓库的更新动态。

- 1CACSurv: Concordance-Aligned Comparative Learning with Large Language Models for Cancer Survival Prediction香港科技大学 · 2026年



