遇见数据集

An m/z- and intensity-based HRMS clustering algorithm targeted towards single-cell metabolomics

收藏
Zenodo2025-11-06 更新2026-05-26 收录
官方服务:

资源简介:

Single-cell (SC) metabolomics holds great potential in the development of novel diagnostic tools and mechanistic insights into cell biology. Using high-resolution mass spectrometry (HRMS), the masses of a single cell’s constituents can be determined with an accuracy high enough to derive their respective elemental compositions. Using a molecule’s mass and its MS fragmentation pattern, in many cases a molecular structure can be assigned or looked up in databases. Due to the small measurement cell volume of an Orbitrap MS instrument, samples are scanned multiple times, which necessitates across-scan clustering per sample, and across-sample alignment of m/z values. However, existing HRMS data processing software is not designed to process SC HRMS data, as it typically requires liquid chromatography retention times or reference spectra for m/z clustering and alignment. Furthermore, both binning and density-based clustering have their disadvantages, leading to both peak aggregation and peak splitting.- Herein, a novel, robust SC HRMS m/z clustering and alignment algorithm is presented and compared with two commercially available and industrially standard algorithms, Sciex MarkerView and Thermo FreeStyle. Furthermore, output is compared with clustering results from DBSCAN and binning. Global Clustering unTargeted Analysis (GCTA) enforces a strict maximum on the cluster size, thereby reducing the chance of peak aggregation, and allows for the preservation of sample information for subsequent compound backtracking. Across-scan clustering and across-sample alignment were contrasted for accuracy in finding peaks identified by commercial software output and peaks with known m/z values corresponding to standards and HMDB and LipidMaps database hits. Comparisons are made based on data recorded for quality control samples containing standard mixes as well as SC HRMS data recorded for two different cell lines. This work shows that the presented algorithm is comparable in accuracy with respect to MarkerView and FreeStyle, reliably identifies compounds, is less prone to peak splitting and successfully filters noise. Furthermore, it is shown to be competitive with DBSCAN, binning and MarkerView when compared to theoretical m/z values based on database hits. Lastly, when compared with binning approaches, GCTA is less sensitive to peak splitting while maintaining similar accuracy in clustering peaks. GCTA encompasses both m/z clustering and reference-free alignment, which makes it pivotal to further development of untargeted SC HRMS metabolomics.

提供机构:
Zenodo
创建时间:
2025-11-06
二维码
社区交流群
二维码
科研交流群
商业服务