Oral Squamous Cell Carcinoma - Mass Spectrometry Imaging
收藏资源简介:
The dataset was first featured in Widlak, Piotr, et al. "Detection of molecular signatures of oral squamous cell carcinoma and normal epithelium–application of a novel methodology for unsupervised segmentation of imaging mass spectrometry data." <em>Proteomics</em> 16.11-12 (2016): 1613-1621. For the tissue sample's biochemical preparation details, please refer to the original publication. The biological material was collected from five patients who underwent surgery due to Oral Squamous Cell Carcinoma (OSCC). Tissue samples contained both tumor and surrounding healthy tissue. Each specimen was cut into 10 µm sections in a cryostat. During the sample preparation for the MS imaging, a high-resolution optical scan of each section was captured. Tissue sections were subjected to peptide imaging with the use of a MALDI ToF mass spectrometer. Spectra were recorded within <em>m/z</em> range of 800-4,000. A raster width of 100 µm was applied, and 400 shots were collected from each ablation point. The obtained dataset consisted of 45,738 raw spectra with 109,568 mass channels. An experienced pathologist analyzed the optical scan obtained during the data acquisition process, and tissue regions were annotated. For the highest confidence of the results obtained in this work, we will focus on the two tissue samples out of the entire dataset (8,005 and 11,869 spectra), which have the highest confidence labels, as explained by the pathologist. The preprocessing of the spectra was conducted in MATLAB. Standard preprocessing steps were applied to the spectra. Spectra were resampled to unify the <em>m/z</em> axis across the dataset. The baseline was removed with MATLAB procedure <em>msbackadj()</em> from the Bioinformatics Toolbox. Peaks were aligned using Fast Fourier Transform-based spectral alignment. The TIC normalization ensured a similar intensity level for all spectra. Finally, a GMM approach was used to model the spectra. GMM locates the peak but also estimates the peak area instead of a raw magnitude provided by most methods. Note that the peaks in MSI spectra are right-skewed, so the neighboring GMM components resulting from that phenomenon were identified and merged to better correspond to actual chemical compounds. The resulting dataset is characterized by 3,714 GMM components corresponding to MSI spectrum peaks.
本数据集首次刊载于Widlak Piotr等人的研究论文《口腔鳞状细胞癌与正常上皮的分子标志物检测——成像质谱数据无监督分割新方法的应用》,发表于*Proteomics* 2016年第16卷第11-12期,页码范围为1613-1621。关于组织样本的生化制备细节,请参阅原始文献。本次研究的生物样本取自5名因口腔鳞状细胞癌(Oral Squamous Cell Carcinoma, OSCC)接受手术治疗的患者。所获组织样本同时包含肿瘤组织与周围健康组织。每份标本在冰冻切片机中被切成10 µm厚度的切片。在进行质谱成像的样本制备过程中,对每一张切片进行了高分辨率光学扫描。组织切片采用基质辅助激光解吸电离飞行时间(MALDI ToF)质谱仪完成肽段成像。质谱信号的采集范围设置为质荷比(m/z)800~4000区间。采用100 µm的光栅步距,且每个消融位点采集400次激光脉冲。所得原始数据集包含45738条原始质谱图,共涵盖109568个质量通道。一名经验丰富的病理学家对数据采集过程中获得的光学扫描图像进行分析,并对组织区域进行标注。为保证本研究结果的最高置信度,我们将聚焦于整个数据集中的两份组织样本(分别对应8005和11869条质谱图),正如病理学家所述,这两份样本拥有最高置信度的标注标签。质谱图的预处理工作在MATLAB环境中完成。我们采用标准预处理步骤对质谱图进行处理:首先将所有质谱图重采样以统一数据集内的质荷比(m/z)轴;借助生物信息学工具箱(Bioinformatics Toolbox)中的MATLAB函数*msbackadj()*去除基线漂移;采用基于快速傅里叶变换的谱对齐算法完成峰对齐;通过总离子流(Total Ion Current, TIC)归一化使所有质谱图的强度水平趋于一致。最后采用高斯混合模型(Gaussian Mixture Model, GMM)对质谱图进行建模。GMM不仅可以定位峰位置,还能估算峰面积,而非多数方法所提供的原始峰强度。需注意的是,质谱成像光谱中的峰呈右偏分布,因此针对该现象产生的相邻GMM分量将被识别并合并,以更好地匹配实际的化学化合物。最终得到的数据集包含3714个GMM分量,对应质谱成像光谱的峰特征。



