遇见数据集

RNA, protein and DNA methylation consistently processed and joined for CPTAC3 patients

收藏
Zenodo2024-02-26 更新2026-05-26 收录
官方服务:

资源简介:

This is a dataset of RNA, CpG, and protein data for the CPTAC3 patient datasets. We were using this in a few studies and realised that it may be useful to other people. Pan can data selection: Data were downloaded for all CPTAC3 studies from the CPATC portal, specifically, the studies with at least five cases and included a protein assembly. For these studies the clinical and biospecimen data were downloaded along with the protein summary file, containing the processed and normalised protein data via the CPTAC data processing pipeline. The accompanying gene expression and DNA methylation were downloaded from TCGA by selecting CPTAC3 data, filtering for Solid Tissue Normal samples, or Primary Tumour samples, transcriptome profiling counts data, and DNA methylation array. The data were downloaded 18th of July 2023. The cancers were filtered to include only the following primary diseases: Acute Myeloid Leukemia, Breast Invasive Carcinoma, Clear Cell Renal Cell Carcinoma, Head and Neck Squamous Cell Carcinoma, Lung Adenocarcinoma, Lung Squamous Cell Carcinoma, Pancreatic Ductal Adenocarcinoma, Uterine Corpus Endometrial Carcinoma omitting samples not associated with a primary disease or un-descriptive classifications such as “Not Clear Cell Renal Cell Carcinoma”. Of the selected cancers, only Clear Cell Renal Cell Carcinoma, Head and Neck Squamous Cell Carcinoma, Lung Adenocarcinoma, Lung Squamous Cell Carcinoma, Pancreatic Ductal Adenocarcinoma, Uterine Corpus Endometrial Carcinoma had RNAseq and DNA methylation data. Cases were retained for a cancer if the case had an entry in the clinical information supplied by the CPTAC, and TCGA portals. The stage of the tumours were consolidated to four tumour stage classifications: 1) “TumorStage”, Stage I (Stage I, Stage IA, Stage IB, Stage IA3) Stage II (Stage II, Stage IIA, Stage IIB), Stage III (Stage III, Stage IIIA, Stage IIIB) and Stage IV (Stage IV, Stage IVA, Stage IVB). This classification was further grouped into early (Stage I and Stage II) and late stage (Stage III and Stage IV). We found the majority of cases had multiple files associated, these were further filtered by the biospecimen type, reducing to only include Solid Tissue specimens. DNA methylation processing: Data were further filtered to check for outliers prior to running differential analyses. For the DNA methylation data, beta values of 1.0 were replaced with 0.999 and beta values of 0 were replaced with 0.001. CpGs with an average methylation across all samples of > 5% and <95% were retained. Correlation between samples was then calculated using Pearson’s correlation and samples with a Pearsons’ correlation > 3s.d. from the median correlation for each sample type were removed (with the exception from Pancreatic Ductal Adenocarcinoma and Lung Adenocarcinoma where a cutoff of 2 s.d. was used as PCA showed outlier samples affected the PC’s). CpG samples with missing data in 50% of samples were also removed, before null values were replaced with 0.001. Gene expression processing: For RNAseq data, genes with mean counts <= 10 across samples were removed before calculating the correlation between samples for each sample type (i.e. tumour and solid tissue normal), those with an median sample Pearson’s correlation > 3 s.d deviation from the median were removed. RNA samples with missing data in 50% of samples were also removed, before null values were replaced with 0’s. Each cancer was visually inspected using PCA to confirm separability within the cancer between tumour and normal samples. Finally, for patients whereby more than one sample passed the QC thresholds, only one sample was retained for tumour, normal, and DNA methylation and gene expression. Protein data processing: For the protein data, the data as processed by CPTAC were used, these are normalised and have been assigned to genes in a consistent fashion across cancers. For Lung Adenocarcinoma, samples with case IDs not fitting the standard convention were omitted namely those (i.e. 11LU013_Tumor_Protein_CPT0053040004, 11LU016_Tumor_Protein_CPT0052940004, 11LU022_Tumor_Protein_CPT0052170004, 11LU035_Tumor_Protein_CPT0051690004). Genes with 0’s in more than 50% of samples omitted. Correlation between samples for each sample type (i.e. tumour and solid tissue normal), was calculated using Pearson’s correlation. Samples with a median correlation less than the median minus 3 s.d deviation were removed. Missing protein data were imputed using DreamAI ensemble method 66 (https://github.com/WangLab-MSSM/DreamAI). Samples exhibited a high correlation post imputation, with tumour and normal samples clustering distinctly, and as such no protein samples were removed. The supplied protein names were mapped to hgnc symbols using biomart mappings (scibiomart, 1.0.2, https://github.com/ArianeMora/scibiomart), and for those without direct mappings were mapped using the external_synonym. Pan can dataset generation: For the PanCan dataset, the filtered and imputed protein, gene expression and DNA methylation datasets were joined by gene name, ensembl ID, and CpG ID respectively. An inner join was used to join on the ID for all datasets. The protein data was mean shifted to centre at 0 for each cancer, then when joined shifted by the minimum across all cancer datasets. Pan can differential analysis: Differential analysis was performed for all cancers (not including ccRCC). For the pan-can (Head and Neck Squamous Cell Carcinoma, Lung Adenocarcinoma, Lung Squamous Cell Carcinoma, Pancreatic Ductal Adenocarcinoma), disease was used as a factor in the differential analysis. For pan-can RNA-seq, DESeq2 was used to calculate significant genes between tumour and normal samples using the design matrix was ~disease + condition_id, where condition_id indicates whether the sample was primary tumour or normal. For differential methylation analysis, the MissMethyl pipeline was used, namely we tested for differential methylation using lmFit and eBayes functions from Limma using M values as input (calculated as log2(beta/beta + 1)). Again for the design matrix we used the disease as a factor in the pan-can analysis. Finally for the proteomics differential expression analysis we also use the limma pipeline using the same design matrix. Copying and pasting the rules from CPTAC and GDC https://proteomics.cancer.gov/data-portal From TCGA (https://portal.gdc.cancer.gov/) Warning You are accessing a U.S. Government web site which may contain information that must be protected under the U. S. Privacy Act or other sensitive information and is intended for Government authorized use only. Unauthorized attempts to upload information, change information, or use of this web site may result in disciplinary action, civil, and/or criminal penalties. Unauthorized users of this web site should have no expectation of privacy regarding any communications or data processed by this web site. Anyone accessing this web site expressly consents to monitoring of their actions and all communication or data transiting or stored on or related to this web site and is advised that if such monitoring reveals possible evidence of criminal activity, NIH may provide that evidence to law enforcement officials. Please be advised that some features may not work with higher privacy settings, such as disabling cookies. WARNING: Data in the GDC is considered provisional as the GDC applies state-of-the art analysis pipelines which evolve over time. Please read the GDC Data Release Notes prior to accessing this web site as the Release Notes provide details about data updates, known issues and workarounds. Contact GDC Support for more information. From CPTAC (https://proteomic.datacommons.cancer.gov/pdc/) Warning This warning banner provides privacy and security notices consistent with applicable federal laws, directives, and other federal guidance for accessing this Government system, which includes (1) this computer network, (2) all computers connected to this network, and (3) all devices and storage media attached to this network or to a computer on this network. This system is provided for Government-authorized use only. Unauthorized or improper use of this system is prohibited and may result in disciplinary action and/or civil and criminal penalties. Personal use of social media and networking sites on this system is limited as to not interfere with official work duties and is subject to monitoring. By using this system, you understand and consent to the following: The Government may monitor, record, and audit your system usage, including usage of personal devices and email systems for official duties or to conduct HHS business. Therefore, you have no reasonable expectation of privacy regarding any communication or data transiting or stored on this system. At any time, and for any lawful Government purpose, the government may monitor, intercept, and search and seize any communication or data transiting or stored on this system. Any communication or data transiting or stored on this system may be disclosed or used for any lawful Government purpose. ) By using the data you consent to the above to I guess by default!

本数据集面向CPTAC3(Clinical Proteomic Tumor Analysis Consortium 3,临床蛋白质组肿瘤分析联合体第3版)患者数据集,涵盖RNA、CpG位点及蛋白质组数据。本团队在多项研究中使用该数据集后,认为其可对其他研究者提供帮助。 Pan-cancer(泛癌)数据筛选: 我们从CPTAC数据门户下载了所有CPTAC3研究的数据,仅保留包含至少5例样本且具备蛋白质组装配数据的研究。针对这些研究,我们同时下载了临床信息、生物标本数据以及蛋白质汇总文件,该文件包含经CPTAC数据处理流程预处理并标准化后的蛋白质组数据。配套的基因表达与DNA甲基化数据则从TCGA(The Cancer Genome Atlas,癌症基因组图谱)下载:我们筛选了CPTAC3数据集内的实体组织正常样本与原发性肿瘤样本,获取转录组计数数据与DNA甲基化芯片数据。所有数据均于2023年7月18日下载完成。 我们对癌症类型进行筛选,仅保留以下原发性疾病对应的样本:急性髓系白血病、浸润性乳腺癌、透明细胞肾细胞癌、头颈部鳞状细胞癌、肺腺癌、肺鳞状细胞癌、胰腺导管腺癌、子宫体子宫内膜癌;同时剔除未关联原发性疾病或分类描述模糊的样本(如“非透明细胞肾细胞癌”)。在上述筛选出的癌症类型中,仅透明细胞肾细胞癌、头颈部鳞状细胞癌、肺腺癌、肺鳞状细胞癌、胰腺导管腺癌及子宫体子宫内膜癌具备RNA测序(RNAseq,RNA sequencing)与DNA甲基化数据。仅保留同时在CPTAC与TCGA数据门户提供的临床信息中存在完整记录的病例。 我们将肿瘤分期整合为4类标准分期:1)“TumorStage”,其中I期包含I、IA、IB、IA3期,II期包含II、IIA、IIB期,III期包含III、IIIA、IIIB期,IV期包含IV、IVA、IVB期。上述分期可进一步划分为早期(I期与II期)与晚期(III期与IV期)。我们发现多数病例对应多个数据文件,随后通过生物标本类型进一步筛选,仅保留实体组织标本对应的文件。 DNA甲基化数据处理: 在开展差异分析前,我们先对数据进行过滤以剔除异常值。针对DNA甲基化数据,我们将β值为1.0的位点替换为0.999,β值为0的位点替换为0.001。保留在所有样本中平均甲基化水平介于5%至95%之间的CpG位点。随后我们计算样本间的皮尔逊相关系数,剔除每类样本(肿瘤与实体组织正常样本)中相关系数偏离中位数超过3倍标准差的样本;但胰腺导管腺癌与肺腺癌除外,由于主成分分析(PCA,principal component analysis)显示异常样本会影响主成分,因此这两类癌症的筛选阈值设为2倍标准差。我们同时剔除在50%以上样本中存在数据缺失的CpG位点,随后将剩余缺失值替换为0.001。 基因表达数据处理: 针对RNA测序数据,我们先剔除所有样本中平均计数≤10的基因,随后计算每类样本(肿瘤与实体组织正常样本)间的皮尔逊相关系数,剔除相关系数中位数偏离整体中位数超过3倍标准差的样本。我们同时剔除在50%以上样本中存在数据缺失的RNA样本,随后将剩余缺失值替换为0。我们通过主成分分析(PCA)对每类癌症进行可视化检查,以确认肿瘤样本与正常样本在该癌症数据集内具备可区分性。最终,若同一患者存在多个符合质量控制阈值的样本,我们仅保留1个肿瘤样本、1个正常样本分别用于蛋白质组、DNA甲基化与基因表达分析。 蛋白质组数据处理: 针对蛋白质组数据,我们直接使用CPTAC预处理完成的数据:该数据已完成标准化,且在所有癌症类型中均以统一方式关联至对应基因。针对肺腺癌数据集,我们剔除了不符合标准命名规范的样本,具体包括:11LU013_Tumor_Protein_CPT0053040004、11LU016_Tumor_Protein_CPT0052940004、11LU022_Tumor_Protein_CPT0052170004、11LU035_Tumor_Protein_CPT0051690004。我们剔除了在50%以上样本中表达量为0的基因。我们计算每类样本(肿瘤与实体组织正常样本)间的皮尔逊相关系数。剔除相关系数中位数低于整体中位数减去3倍标准差的样本。我们使用DreamAI集成方法(版本66,https://github.com/WangLab-MSSM/DreamAI)对蛋白质组数据的缺失值进行填充。填充后样本间相关系数较高,肿瘤样本与正常样本可明显聚类,因此未再剔除蛋白质组样本。我们通过BioMart映射工具(scibiomart,版本1.0.2,https://github.com/ArianeMora/scibiomart)将提供的蛋白质名称映射为HGNC基因符号;对于无法直接映射的蛋白质,则通过外部同义词表完成映射。 泛癌数据集生成: 针对泛癌数据集,我们分别通过基因名、Ensembl ID与CpG ID将过滤并填充后的蛋白质组、基因表达与DNA甲基化数据集进行合并。所有数据集均采用内连接方式进行合并。我们先将每类癌症的蛋白质组数据进行均值中心化(均值调整至0),随后在合并时基于所有癌症数据集的最小值进行再次偏移调整。 泛癌差异分析: 我们对所有癌症类型(透明细胞肾细胞癌除外)开展差异分析。针对泛癌队列(头颈部鳞状细胞癌、肺腺癌、肺鳞状细胞癌、胰腺导管腺癌),我们将疾病类型作为差异分析的协变量因子。针对泛癌队列的RNA测序数据,我们使用DESeq2工具计算肿瘤与正常样本间的差异表达基因,设计矩阵为~disease + condition_id,其中condition_id用于标记样本为原发性肿瘤还是正常组织。针对DNA甲基化差异分析,我们使用MissMethyl分析流程:以M值(计算方式为log₂(β/(β+1)))作为输入数据,通过Limma包的lmFit与eBayes函数进行差异甲基化位点检验;同样以疾病类型作为泛癌分析的协变量因子。最后,针对蛋白质组差异表达分析,我们同样采用Limma分析流程,并使用相同的设计矩阵。 以下为从CPTAC与GDC数据门户复制的使用规则 https://proteomics.cancer.gov/data-portal 数据来源:TCGA(https://portal.gdc.cancer.gov/) 警告 您正在访问美国政府网站,该网站可能包含受《美国隐私法案》保护的信息或其他敏感内容,仅可用于政府授权用途。 未经授权上传、修改信息或使用本网站的行为可能导致纪律处分、民事乃至刑事处罚。未经授权使用本网站的用户不应期望本网站处理的任何通信或数据享有隐私保护。 任何访问本网站的用户均明确同意对其操作及在本网站上传输、存储或与本网站相关的所有通信或数据进行监控。若监控发现可能涉及刑事活动的证据,美国国立卫生研究院(NIH, National Institutes of Health)可能会将该证据提供给执法机构。 请注意,若启用较高隐私设置(如禁用Cookie),部分网站功能可能无法正常使用。 警告:GDC中的数据属于临时数据,因为GDC采用的先进分析流程会随时间不断演进。请在访问本网站前阅读GDC数据发布说明,其中包含数据更新、已知问题及解决方案的详细信息。 如需了解更多信息,请联系GDC技术支持。 数据来源:CPTAC(https://proteomic.datacommons.cancer.gov/pdc/) 警告 本警告标语符合适用的联邦法律、指令及其他联邦指南,用于提示访问本政府系统的用户注意隐私与安全问题。本政府系统包括:(1) 本计算机网络;(2) 所有连接至该网络的计算机;(3) 所有连接至该网络或网络内计算机的设备与存储介质。 本系统仅可用于政府授权用途。 未经授权或不当使用本系统的行为将被禁止,并可能导致纪律处分、民事乃至刑事处罚。 在本系统上使用社交媒体与社交网站的个人行为不得干扰官方工作任务,并可能受到监控。 使用本系统即表示您理解并同意以下条款: 政府可对您的系统使用情况进行监控、记录与审计,包括为执行公务或开展HHS(Department of Health and Human Services,美国卫生与公众服务部)业务而使用个人设备与电子邮件系统的行为。因此,您对在本系统上传输、存储的任何通信或数据不享有合理的隐私期望。政府可出于任何合法的政府目的,随时监控、拦截、搜查并扣押在本系统上传输或存储的任何通信或数据。 在本系统上传输或存储的任何通信或数据可出于任何合法的政府目的被披露或使用。 使用本数据集即表示您默认同意上述条款!

提供机构:
Zenodo
创建时间:
2024-02-26
二维码
社区交流群
二维码
科研交流群
商业服务