遇见数据集

Oral Squamous Cell Carcinoma - Mass Spectrometry Imaging

收藏
NIAID Data Ecosystem2026-05-02 收录
数据链接:
官方服务:

资源简介:

The dataset was first featured in Widlak, Piotr, et al. "Detection of molecular signatures of oral squamous cell carcinoma and normal epithelium–application of a novel methodology for unsupervised segmentation of imaging mass spectrometry data." Proteomics 16.11-12 (2016): 1613-1621. For the tissue sample's biochemical preparation details, please refer to the original publication. The biological material was collected from five patients who underwent surgery due to Oral Squamous Cell Carcinoma (OSCC). Tissue samples contained both tumor and surrounding healthy tissue. Each specimen was cut into 10 µm sections in a cryostat. During the sample preparation for the MS imaging, a high-resolution optical scan of each section was captured. Tissue sections were subjected to peptide imaging with the use of a MALDI ToF mass spectrometer. Spectra were recorded within m/z range of 800-4,000. A raster width of 100 µm was applied, and 400 shots were collected from each ablation point. The obtained dataset consisted of 45,738 raw spectra with 109,568 mass channels. An experienced pathologist analyzed the optical scan obtained during the data acquisition process, and tissue regions were annotated. For the highest confidence of the results obtained in this work, we will focus on the two tissue samples out of the entire dataset (8,005 and 11,869 spectra), which have the highest confidence labels, as explained by the pathologist. The preprocessing of the spectra was conducted in MATLAB. Standard preprocessing steps were applied to the spectra. Spectra were resampled to unify the m/z axis across the dataset. The baseline was removed with MATLAB procedure msbackadj() from the Bioinformatics Toolbox. Peaks were aligned using Fast Fourier Transform-based spectral alignment. The TIC normalization ensured a similar intensity level for all spectra. Finally, a GMM approach was used to model the spectra. GMM locates the peak but also estimates the peak area instead of a raw magnitude provided by most methods. Note that the peaks in MSI spectra are right-skewed, so the neighboring GMM components resulting from that phenomenon were identified and merged to better correspond to actual chemical compounds. The resulting dataset is characterized by 3,714 GMM components corresponding to MSI spectrum peaks.

本数据集首次见于Widlak、Piotr等人发表于《Proteomics》2016年第16卷第11-12期的论文《口腔鳞状细胞癌与正常上皮的分子特征检测——成像质谱数据无监督分割新方法的应用》(Detection of molecular signatures of oral squamous cell carcinoma and normal epithelium–application of a novel methodology for unsupervised segmentation of imaging mass spectrometry data),页码为1613-1621。组织样本的生化制备细节请参见原始文献。 生物样本取自5名因口腔鳞状细胞癌(Oral Squamous Cell Carcinoma, OSCC)接受手术治疗的患者,组织样本同时包含肿瘤组织与周边正常健康组织。 所有标本均在冰冻切片机中被切成10 µm厚度的切片。在质谱成像(mass spectrometry imaging, MSI)样本制备过程中,对每一张切片进行了高分辨率光学扫描。 组织切片采用基质辅助激光解吸电离飞行时间质谱(matrix-assisted laser desorption/ionization time-of-flight mass spectrometer, MALDI-TOF)进行肽段成像。质谱信号的采集范围设置为质荷比(mass-to-charge ratio, m/z)800~4000,采用100 µm的步径进行扫描,并在每个消融点采集400次激光轰击信号。最终得到的数据集包含45738条原始质谱,共109568个质量通道。 一名经验丰富的病理学家对数据采集过程中获得的光学扫描图像进行分析,并对组织区域进行标注。为保障本研究结果的最高置信度,我们将聚焦于整个数据集中标注置信度最高的两份组织样本(分别对应8005条和11869条质谱信号),如病理学家所述。 质谱信号的预处理工作在MATLAB环境中完成,采用标准预处理流程对所有质谱进行处理。首先对质谱进行重采样,以统一全数据集的质荷比轴;随后通过MATLAB生物信息学工具箱(Bioinformatics Toolbox)中的msbackadj()函数去除基线噪音;基于傅里叶变换的光谱对齐算法完成峰位对齐;采用总离子流(Total Ion Current, TIC)归一化方法使所有质谱的信号强度处于相近水平。最后使用高斯混合模型(Gaussian Mixture Model, GMM)对质谱进行建模:高斯混合模型不仅可以定位峰位,还能估算峰面积,而非多数方法所采用的原始峰强度。需注意,质谱成像光谱中的峰呈现右偏分布,因此针对该现象产生的相邻高斯混合模型分量进行识别与合并,以更好地匹配实际的化学化合物。最终得到的数据集包含3714个对应质谱成像光谱峰的高斯混合模型分量。

创建时间:
2024-07-15
二维码
社区交流群
二维码
科研交流群
商业服务