遇见数据集

Malignant Peritoneal Mesothelioma: A phase II clinical trial data

收藏
Zenodo2025-10-06 更新2026-05-26 收录
官方服务:

资源简介:

Malignant Peritoneal Mesothelioma (MPM) flow cytometry dataset is from the paper “Adjuvant dendritic cell-based immunotherapy after cytoreductive surgery and hyperthermic intraperitoneal chemotherapy in patients with malignant peritoneal mesothelioma- a phase II clinical trial”. The clinical trial contains samples from 14 patients collected at 3 time points during their immunotherapy – baseline (before starting vaccination), 2 weeks after the first vaccine, and 2 weeks after the third vaccine. The 14 patients were divided into 2 groups (Group 1 with 9 patients and Group 2 with 5 patients) for sample collection. The collected PBMC samples were analyzed under 6 panels. We have preprocessed the samples under the 3 panels that were targeted to analyze the T-cell populations. The names of these panels are – Co-inhibition, Co-stimulation, and Cytokine. There is another paper “CytoNorm 2.0- A flexible normalization framework for cytometry data without requiring dedicated controls” that worked on 2 other panels of the MPM dataset targeted to analyze B-cell populations. We received the raw data from the authors of the MPM paper in .fcs format and preprocessed them following the steps mentioned in their paper’s supplementary materials. The preprocessing steps involved margin event removal, compensation, transformation, scaling, quality control, manual gating for Lymphocytes and normalization. We have used PeacoQC, flowCore and CytoNorm2.0 for these steps. These are all Bioconductor packages available in R. After preprocessing, we have stored the data in .xlsx format. Under the Final Dataset folder, you will find separate sub-folders for the 3 T-cell panels. Each of the panel folders contain 3 sub-folders for 3 time points. Each time point folder has the preprocessed data files in .xlsx format. The file naming convention of the original data files was quite confusing, so we have renamed every file for convenience of analysis. The naming convention is as follows: NormalizationStep_PanelName_GroupID_TubeID_PatientCode_TimePoint_QualityControlled_ParentPopulation_CellClassification.xlsx. For example, file with the name Norm_Norm_coinhib_g1_Tube_001_MCV001_VAC_1_QC_Lymph_CellType.xlsx means normalization was done on the original data in 2 steps, the file comes from Co-inhibition panel, the patient was in Group 1, the sample tube was Tube_001, patient code was MCV001, time point was 1, the sample is quality controlled, the parent cell population is Lymphocytes, and the file contains cell type label. Note that, for the time point information, VAC_1 means baseline, VAC_2 means after first vaccine, and VAC_3 means after third vaccine. For all samples, the cell classification (cell type assignment) was done using a semi-automated gating procedure guided by limited manual gating by a domain expert. The original clinical study started with 16 patients, but patient MCV008 and MCV011 did not complete the trial. The metadata.xlsx file has the original and the new file names for the record. This file also has information about the overall survival (OS) and progression free survival (PFS) of the patients (in months). The data files are in tabular format. Each row is an individual cell, and each column is a feature (except the last column which is the cell type label). The features include forward and side scatter information (Time, FSC and SSC values) and the cell markers. Do not confuse this Time feature with the 3 time points of sample collection, this Time parameter was recorded while the flow cytometer ran and might not be required further because the data preprocessing is already done. The actual fluorochromes corresponding to the cell markers are also recorded in the file channel.xlsx. The 3 T-cell panels have a common set of cell markers, called the cell type markers - CD56, CD3, CD4, FoxP3, CD8, CD45RA, CCR7, and Live/Dead (the viability channel). Additionally, they have specific cell state markers: LAG3, PD1, TIM3, CD39, KI67, CTLA-4 (Co-inhibition); CD28, CD137, PD1, HLA-DR, ICOS, KI67 (Co-stimulation); PD1, TBET, IL-10, TNF-a, IL2, IFN-y (Cytokine). Note that, the column ordering of the features in the data files may not always be the same. For unsupervised analysis you can exclude the cell type label column. For supervised analysis, you can use our cell type labeling. As this is not a universal labeling, one can use their choice of clustering/ classification algorithm to classify the cells as well. Based on the common cell type markers, we identified the following cell populations: NK cell (natural killer), NK T-cell, CD4+ Treg (CD4+ regulatory T-cell), CD4+ EM (CD4+ effector memory T-cell), CD4+ EMRA (CD4+ terminally differentiated effector memory T-cell), CD4+ CM (CD4+ central memory T-cell), CD4+ Naïve T-cell, CD8+ EM, CD8+ EMRA, CD8+ CM, CD8+ Naïve T-cell, Other CD3+ (uncategorized CD3+ cell), Other CD3- (uncategorized CD3- cell). If there is any outlier cell that does not fall under any of the above populations, it is labeled as “Other cell” and should be excluded in the analysis. The Raw Data folder contains the files in .fcs format along the preporocesing steps.

提供机构:
Zenodo
创建时间:
2025-10-06
二维码
社区交流群
二维码
科研交流群
商业服务