遇见数据集

A unified benchmark of synthetic data generation for clinical transcriptomic cancer cohorts

收藏
Zenodo2026-05-01 更新2026-05-26 收录
官方服务:

资源简介:

Achieving a trade-off between biological utility and patient privacy remains a key challenge for secure data sharing when applying transcriptomic clinical datasets to artificial intelligence in precision oncology. Here, we introduce the first benchmarking study tailored to high-dimensional clinical transcriptomic cancer data, comparing synthetic data generation methods across three clinical cancer trials. Our framework, SynOmicBench, combines standardized preprocessing with multidimensional evaluation, prioritizing downstream biological validation alongside statistical fidelity and attack-based privacy assessment. Results indicate that no single method dominated all dimensions, with Gaussian Copula achieving the most balanced performance, followed by Avatar, demonstrating that metric-based similarity alone is insufficient to ensure preservation of higher-order molecular dependencies. Synthetic data consistently reproduced biomedical signal directionality but with attenuated effect sizes and inter-replicate variability, supporting hypothesis generation when multi-seed synthesis is adopted. Collectively, this framework provides a reproducible decision-support tool for method selection and promotes biologically informed, privacy-aware adoption of synthetic data in precision oncology.

在精准肿瘤学领域将转录组学临床数据集应用于人工智能时,实现生物学效用与患者隐私之间的平衡,仍是安全数据共享面临的核心挑战。本研究首次提出针对高维临床癌症转录组学数据的基准测试研究,对三项临床癌症试验中的合成数据生成方法开展对比分析。我们构建的SynOmicBench框架将标准化预处理与多维度评估相结合,评估过程中优先考量下游生物学验证、统计保真度以及基于攻击的隐私性评估。研究结果显示,没有任何一种方法能够在所有维度上占据全面优势,其中高斯Copula(Gaussian Copula)的综合性能最为均衡,其次为Avatar,这表明仅依靠基于指标的相似性不足以确保高阶分子依赖关系的保留。合成数据能够稳定复现生物医学信号的方向性,但效应量与重复间变异度有所衰减;当采用多种子合成策略时,该数据集可有效辅助假设生成。总体而言,本框架可为合成数据方法的选择提供可复现的决策支持工具,并推动精准肿瘤学领域中基于生物学考量、兼顾隐私保护的合成数据应用。

提供机构:
Zenodo
创建时间:
2026-04-20
二维码
社区交流群
二维码
科研交流群
商业服务