Single Cell CPTAC Renal Cell Carcinoma
收藏资源简介:
These data include a subset of single-cell samples from the CPTAC Renal Cell Carcinoma data processed using the following steps: Loom files were read into R and converted into SingleCellExperiment objects. Ensembl gene ID's were matched to HGNC symbols, chromosome name, starting position, and ending position via Biomart using the scater R package. Size factors were computed using the scran R package. UMAP dimensions were computed using the scater R package. Probes without matching HGNC symbols were removed. Where duplicate HGNC symbols were present, the gene with the maximum normalized range was retained. Cell types were inferred using the scMRMA R package. Cells with less than or equal to 1,000 features were removed. Cells with mitochondrial reads greater than or equal to 50% were removed. Expression data for podocytes and macrophages were saved separately for each sample. For podocytes and macrophages for each sample, expression data were projected onto the first 100 principal components using the irlba R package. The results are available from CPTAC_RCC_PCA.zip. For macrophages for each sample, a differentiation trajectory was estimated using the monocle3 R package, and plots were colored by combined expression of the M0 markers CSF1R, CD14, CD68, and CD11B, the M1 markers CD86, MARC0, CXCL9, CXCL10, CXCL11, NOS2, SOCS1, and CD64, and the M2 markers TGM2, CD23, ARG1, CCL22, CD163, and CD206 (from PMC8268869). Pseudotime starting points were annotated in monocle3 using visual inspection of plots. Only samples in which a visible trajectory from M0 -> M1 -> M2 was evident using these markers were retained. For podocytes for each sample, a differentiation trajectory was estimated using the monocle3 R package, and plots were colored by combined expression of the dedifferentiation markers DACH1 (from PMC5908116) and PTPRO (from PMID9639039. Pseudotime starting points were annotated in monocle3 using visual inspection of plots. Only samples in which a visible trajectory of dedifferentiation was evident using these markers were retained. These pseudotime assignments and the macrophage assignments are available from CPTAC_RCC_pseudotime_all_cells.zip. For podocytes and macrophages for each sample, 10-fold cross-validation matrices for expression, PCA, and pseudotime were generated across 5 random splits, for a total of 50 files per cell type, per sample. These files are available from CPTAC_RCC_expression_all_cells.zip. Expression data were subset to include 63 randomly-selected podocytes and 63 randomly-selected macrophages to ensure balanced data. These expression data are available from CPTAC_RCC_expression.zip. Pseudotimes were also subset and are available in CPTAC_RCC_pseudotime.zip. Expression data were projected onto the first 100 principal components. These data are available from CPTAC_RCC_expression_dimReduced.zip.
本数据集包含取自临床蛋白质组肿瘤分析联盟(Clinical Proteomic Tumor Analysis Consortium, CPTAC)肾细胞癌(Renal Cell Carcinoma, RCC)数据集的单细胞样本子集,其预处理流程如下: 1. 将Loom文件读取至R语言环境中,并转换为SingleCellExperiment对象。 2. 借助scater R包,通过Biomart数据库将Ensembl基因ID匹配至HGNC基因符号、染色体名称、基因起始位置与终止位置。 3. 使用scran R包计算大小因子(size factors)。 4. 通过scater R包计算UMAP降维维度。 5. 移除未匹配到HGNC符号的探针。 6. 若存在重复的HGNC符号,则保留归一化范围最大的基因。 7. 使用scMRMA R包推断细胞类型。 8. 移除特征数量≤1000的细胞。 9. 移除线粒体reads(mitochondrial reads)占比≥50%的细胞。 10. 将每个样本的足细胞(podocytes)与巨噬细胞(macrophages)的表达数据分别进行保存。 11. 针对每个样本的足细胞与巨噬细胞,使用irlba R包将其表达数据投影至前100个主成分(principal components),结果存储于CPTAC_RCC_PCA.zip文件中。 12. 针对每个样本的巨噬细胞,使用monocle3 R包估算其分化轨迹,并以M0标志物CSF1R、CD14、CD68、CD11B,M1标志物CD86、MARC0、CXCL9、CXCL10、CXCL11、NOS2、SOCS1、CD64,以及M2标志物TGM2、CD23、ARG1、CCL22、CD163、CD206(来源:PMC8268869)的联合表达水平对可视化绘图进行着色。使用monocle3通过目视检查绘图结果来注释拟时间(pseudotime)的起始点,仅保留可通过上述标志物观察到M0→M1→M2分化轨迹的样本。 13. 针对每个样本的足细胞,使用monocle3 R包估算其去分化轨迹,并以去分化标志物DACH1(来源:PMC5908116)与PTPRO(来源:PMID9639039)的联合表达水平对可视化绘图进行着色。使用monocle3通过目视检查绘图结果来注释拟时间起始点,仅保留可通过上述标志物观察到足细胞去分化轨迹的样本。上述拟时间分配结果与巨噬细胞分类结果均存储于CPTAC_RCC_pseudotime_all_cells.zip文件中。 14. 针对每个样本的足细胞与巨噬细胞,通过5次随机划分生成表达数据、PCA结果及拟时间的10折交叉验证矩阵,每种细胞类型对应每个样本共生成50个文件,相关文件存储于CPTAC_RCC_expression_all_cells.zip中。 15. 对表达数据进行随机抽样,分别选取63个足细胞与63个巨噬细胞以保证数据集平衡,该平衡后的表达数据存储于CPTAC_RCC_expression.zip中。对应的拟时间数据也经过相同抽样处理,存储于CPTAC_RCC_pseudotime.zip中。 16. 将表达数据投影至前100个主成分,相关数据存储于CPTAC_RCC_expression_dimReduced.zip文件中。



