Data from: Patient-derived organoids from metastatic colorectal cancer mirror tumor heterogeneity and predict patient survival and drug sensitivity
收藏资源简介:
Datasets Data S1 to S5 are generated in the study “Patient-derived organoids from metastatic colorectal cancer mirror tumor heterogeneity and predict patient survival and drug sensitivity”. The data originate from patient-derived organoids (PDOs) and corresponding colorectal liver metastases (CRLMs) and include genomic, transcriptomic, proteomic, and pharmacological measurements. All datasets are linked through a common sample_id, enabling cross-referencing between molecular, phenotypic, and drug response data from the same samples. Data S1 – Somatic mutation profiles Targeted sequencing results from a custom 20-gene panel in PDOs and matched tumor tissues. Each row represents a detected variant. Columns include: sample_id, patient, sample_type, chromosome, position, ref, alt, n_ref_count, n_alt_count, t_ref_count, t_alt_count, maf (mutant allele fraction), filter, gene, func, syn, and AAChange (variant annotation including transcript, exon, nucleotide and amino acid change). Data S2 – RNA sequencing gene expression counts Gene-level RNA sequencing count matrix of PDO samples. Rows correspond to genes and columns correspond to PDO sample_id. Columns include: ensembl_gene_id, entrezgene_id, hgnc_symbol, followed by columns representing individual PDO samples containing raw read counts mapped to each gene. Data S3 – Protein expression from multiplex immunohistochemistry Quantitative protein expression measurements obtained from multiplex immunohistochemistry (mIHC) analysis of PDO samples. Each row represents one protein measurement in a given PDO sample. Columns include: sample_id, Prot_marker, mIHC_stain_no (multiplex panel identifier), AB_order (antibody staining order), Fluor (Opal fluorophore), mean_express_PDO (mean fluorescence intensity), area_pixels (analyzed image area), and Slide (slide identifier). Data S4 – Raw drug sensitivity screening measurements Primary measurements from medium-throughput drug screening assays performed on PDOs. Each row corresponds to one well measurement in a drug screening plate. Columns include: sample_id, sample_id_drums, run_id, assay_no, library_id, compound_name, compound_fimm, compound_drums, compound_type, concentration, concentration_unit, drow, dcol, plate, signal (raw luminescence), and viability (normalized cell viability). Data S5 – Drug sensitivity scores (DSS) Summary matrix of drug response profiles derived from the raw screening measurements in Data S4. Rows correspond to PDO samples and columns correspond to tested compounds and run_ids. Drug names without suffix were caclulated based on the runs indicated in the run_id_harmonize and drugs with suffix _lib1 or _lib2 indicate the drug library and run_id in which the compound was tested for the respective PDO sample. The values represent Drug Sensitivity Scores (DSS), a quantitative metric calculated from dose–response curves that integrates both drug potency and efficacy to summarize the overall sensitivity of each PDO to a given compound.
数据集S1至S5均来自研究《转移性结直肠癌患者来源类器官可模拟肿瘤异质性并预测患者生存与药物敏感性》(Patient-derived organoids from metastatic colorectal cancer mirror tumor heterogeneity and predict patient survival and drug sensitivity)。本数据集源自患者来源类器官(Patient-derived organoids, PDOs)及其对应的结直肠肝转移(colorectal liver metastases, CRLMs)样本,涵盖基因组学、转录组学、蛋白质组学及药理学检测数据。所有数据集通过统一的sample_id(样本ID)进行关联,可实现同一样本的分子、表型及药物反应数据的交叉对照。 数据集S1——体细胞突变谱(Somatic mutation profiles) 包含PDOs及配对肿瘤组织的定制化20基因面板靶向测序结果。每一行代表一个检测到的变异。列信息包括:sample_id(样本ID)、patient(患者)、sample_type(样本类型)、chromosome(染色体)、position(位置)、ref(参考等位基因)、alt(变异等位基因)、n_ref_count(正常组织参考等位基因计数)、n_alt_count(正常组织变异等位基因计数)、t_ref_count(肿瘤组织参考等位基因计数)、t_alt_count(肿瘤组织变异等位基因计数)、maf(mutant allele fraction,突变等位基因频率)、filter(过滤条件)、gene(基因)、func(功能)、syn(同义性),以及AAChange(包含转录本、外显子、核苷酸及氨基酸变化的变异注释)。 数据集S2——RNA测序基因表达计数(RNA sequencing gene expression counts) 为PDO样本的基因水平RNA测序计数矩阵。行对应基因,列对应PDO样本的sample_id。列信息包括:ensembl_gene_id(Ensembl基因ID)、entrezgene_id(Entrez基因ID)、hgnc_symbol(HGNC基因符号),后续列代表单个PDO样本,包含映射至各基因的原始读段计数。 数据集S3——多重免疫组化蛋白质表达(Protein expression from multiplex immunohistochemistry) 包含通过PDO样本的多重免疫组化(multiplex immunohistochemistry, mIHC)分析获得的定量蛋白质表达检测数据。每一行代表某一PDO样本中的一个蛋白质检测结果。列信息包括:sample_id(样本ID)、Prot_marker(蛋白质标记物)、mIHC_stain_no(multiplex panel identifier,多重染色组标识符)、AB_order(antibody staining order,抗体染色顺序)、Fluor(Opal fluorophore,奥帕尔荧光团)、mean_express_PDO(PDO平均荧光强度)、area_pixels(分析图像像素面积)、Slide(玻片标识符)。 数据集S4——原始药物敏感性筛选检测数据(Raw drug sensitivity screening measurements) 包含在PDOs上开展的中通量药物筛选实验的原始检测结果。每一行对应药物筛选板中的一个孔位检测结果。列信息包括:sample_id(样本ID)、sample_id_drums、run_id(实验运行ID)、assay_no(实验编号)、library_id(文库ID)、compound_name(化合物名称)、compound_fimm、compound_drums、compound_type(化合物类型)、concentration(浓度)、concentration_unit(浓度单位)、drow、dcol、plate(板标识符)、signal(原始发光信号)、viability(归一化细胞存活率)。 数据集S5——药物敏感性评分(Drug Sensitivity Scores, DSS) 为基于数据集S4中原始筛选检测数据推导得到的药物反应谱汇总矩阵。行对应PDO样本,列对应受试化合物及run_id(实验运行ID)。无后缀的药物名称基于run_id_harmonize中指定的实验计算得到,带_lib1或_lib2后缀的药物分别代表该化合物在对应文库及实验运行中针对特定PDO样本完成的检测。矩阵数值为药物敏感性评分(Drug Sensitivity Scores, DSS),这是一种由剂量-反应曲线计算得到的定量指标,可整合药物效能与效力,以概括每个PDO对特定化合物的整体敏感性。



