xpertsystems/hconc011-sample
收藏资源简介:
HC-ONC-011是一个多癌症肿瘤进展与生存队列的样本数据集,来自XpertSystems.ai合成数据工厂的肿瘤学垂直领域。这是一个完全合成的泛癌症队列,涵盖10种癌症类型:非小细胞肺癌(NSCLC)、结直肠癌、乳腺癌、胰腺癌、卵巢癌、肝细胞癌(HCC)、前列腺癌、胶质母细胞瘤(GBM)、黑色素瘤和膀胱癌。数据集具有SEER锚定的发病率分布、按癌症类型分期的阶段分布、全面的癌症特异性生物标志物面板(如NSCLC中的EGFR/ALK/ROS1等)、基于生物标志物和癌症分期的现代治疗方案、肿瘤动力学(指数增长和治疗反应)、ctDNA纵向轨迹(基线/3个月/6个月/12个月的VAF)、RECIST 1.1影像评估(3/6/12/24个月)、多器官转移倾向(肝/肺/脑/骨/腹膜)以及基于17项里程碑试验、TCGA泛癌症图谱和SEER 2023校准的OS/PFS终点。数据集设计为可直接用于泛癌症分析、治疗模式建模、多癌症生存基准测试和生物标志物分层建模,同时保持100%合成,无真实患者数据、无PHI、无重新识别风险。样本包含500名患者和73列数据,格式为CSV单表。
HC-ONC-011 is a synthetic cohort dataset for multi-cancer tumor progression and survival, sourced from the oncology vertical of XpertSystems.ai’s Synthetic Data Factory. This is a fully synthetic pan-cancer cohort covering 10 cancer types: non-small cell lung cancer (NSCLC), colorectal cancer, breast cancer, pancreatic cancer, ovarian cancer, hepatocellular carcinoma (HCC), prostate cancer, glioblastoma (GBM), melanoma, and bladder cancer. The dataset features SEER-calibrated incidence distributions, stage distributions stratified by cancer type, comprehensive cancer-specific biomarker panels (e.g., EGFR/ALK/ROS1 in NSCLC), modern treatment regimens stratified by biomarkers and cancer stage, tumor dynamics (exponential growth and treatment response), longitudinal circulating tumor DNA (ctDNA) trajectories (variant allele frequency, VAF, at baseline, 3-month, 6-month, and 12-month time points), RECIST 1.1 imaging assessments at 3, 6, 12, and 24 months, multi-organ metastatic propensity (liver, lung, brain, bone, peritoneum), and overall survival (OS) and progression-free survival (PFS) endpoints calibrated against 17 landmark clinical trials, the TCGA Pan-Cancer Atlas, and SEER 2023 data. This dataset is designed for direct application in pan-cancer analysis, treatment pattern modeling, multi-cancer survival benchmarking, and biomarker-stratified modeling, while remaining 100% synthetic with no real patient data, no protected health information (PHI), and no re-identification risks. It contains 500 patient samples and 73 data columns, formatted as a single CSV table.




