遇见数据集

Integrating Multidimensional Data for Clustering Analysis With Applications to Cancer Patient Data

收藏
DataCite Commons2025-04-01 更新2024-07-28 收录
官方服务:

资源简介:

Advances in high-throughput genomic technologies coupled with large-scale studies including The Cancer Genome Atlas (TCGA) project have generated rich resources of diverse types of omics data to better understand cancer etiology and treatment responses. Clustering patients into subtypes with similar disease etiologies and/or treatment responses using multiple omics data types has the potential to improve the precision of clustering than using a single data type. However, in practice, patient clustering is still mostly based on a single type of omics data or ad hoc integration of clustering results from individual data types, leading to potential loss of information. By treating each omics data type as a different informative representation from patients, we propose a novel multi-view spectral clustering framework to integrate different omics data types measured from the same subject. We learn the weight of each data type as well as a similarity measure between patients via a nonconvex optimization framework. We solve the proposed nonconvex problem iteratively using the ADMM algorithm and show the convergence of the algorithm. The accuracy and robustness of the proposed clustering method is studied both in theory and through various synthetic data. When our method is applied to the TCGA data, the patient clusters inferred by our method show more significant differences in survival times between clusters than those inferred from existing clustering methods. Supplementary materials for this article, including a standardized description of the materials available for reproducing the work, are available as an online supplement.

高通量基因组技术的进步,结合包括癌症基因组图谱(The Cancer Genome Atlas, TCGA)项目在内的大规模研究,已生成了丰富多样的组学数据(omics data)资源,助力我们更深入地理解癌症病因与治疗应答。利用多组学数据将患者划分为具有相似疾病病因或治疗应答的亚型,相较仅使用单一数据类型的聚类方法,有望提升聚类精度。然而在实际应用中,患者聚类大多仍基于单一组学数据,或是对各单一数据类型的聚类结果进行临时整合,这可能造成信息丢失。 我们将每种组学数据类型视为来自患者的不同信息表征,提出了一种新颖的多视图谱聚类(spectral clustering)框架,以整合来自同一受试者的不同组学数据类型。我们通过非凸优化(nonconvex optimization)框架,学习每种数据类型的权重以及患者间的相似性度量。我们采用交替方向乘子法(Alternating Direction Method of Multipliers, ADMM)迭代求解所提出的非凸问题,并证明了该算法的收敛性。 我们从理论层面与各类合成数据集实验两个维度,对所提出的聚类方法的准确性与鲁棒性展开了研究。将该方法应用于TCGA数据时,我们的方法得到的患者聚类结果,相较于现有聚类方法得到的结果,不同聚类间的生存时间差异更为显著。本文的补充材料(包含可复现本研究所需材料的标准化描述)可作为在线补充资源获取。

提供机构:
Taylor & Francis
创建时间:
2021-09-22
二维码
社区交流群
二维码
科研交流群
商业服务