Integrated Strategy for Unknown EI–MS Identification Using Quality Control Calibration Curve, Multivariate Analysis, EI–MS Spectral Database, and Retention Index Prediction
收藏资源简介:
Compound identification using unknown electron ionization (EI) mass spectra in gas chromatography coupled with mass spectrometry (GC–MS) is challenging in untargeted metabolomics, natural product chemistry, or exposome research. While the total count of EI–MS records included in publicly or commercially available databases is over 900 000, efficient use of this huge database has not been achieved in metabolomics. Therefore, we proposed a “four-step” strategy for the identification of biologically significant metabolites using an integrated cheminformatics approach: (i) quality control calibration curve to reduce background noise, (ii) variable selection by hypothesis testing in principal component analysis for the efficient selection of target peaks, (iii) searching the EI–MS spectral database, and (iv) retention index (RI) filtering in combination with RI predictions. In this study, the new MS-FINDER spectral search engine was developed and utilized for searching EI–MS databases using mass spectral similarity with the evaluation of false discovery rate. Moreover, in silico derivatization software, MetaboloDerivatizer, was developed to calculate the chemical properties of derivative compounds, and all retention indexes in EI–MS databases were predicted using a simple mathematical model. The strategy was showcased in the identification of three novel metabolites (butane-1,2,3-triol, 3-deoxyglucosone, and palatinitol) in Chinese medicine Senkyu for quality assessment, as validated using authentic standard compounds. All tools and curated public EI–MS databases are freely available in the ‘Computational MS-based metabolomics’ section of the RIKEN PRIMe Web site (http://prime.psc.riken.jp).
在非靶向代谢组学、天然产物化学或暴露组学研究中,利用未知电子电离(electron ionization, EI)质谱数据开展气相色谱-质谱联用(gas chromatography coupled with mass spectrometry, GC–MS)的化合物鉴定极具挑战性。尽管公开或商用数据库收录的EI质谱记录总量已超过90万条,但代谢组学领域尚未实现对这一海量数据库的高效利用。为此,我们提出了一种基于整合化学信息学(cheminformatics)方法的“四步”策略,用于鉴定具有生物学意义的代谢物:(1) 构建质量控制校准曲线以降低背景噪声;(2) 借助主成分分析中的假设检验开展变量筛选,实现目标峰的高效选取;(3) 进行EI质谱谱库检索;(4) 结合保留指数(retention index, RI)预测开展保留指数过滤。本研究开发并应用了新型MS-FINDER谱库检索引擎,通过质谱相似度匹配开展EI质谱数据库检索,并对假发现率(false discovery rate)进行评估。此外,本研究还开发了虚拟衍生化(in silico derivatization)软件MetaboloDerivatizer,用于计算衍生化合物的化学性质;同时通过简单数学模型预测了EI质谱数据库中所有保留指数。本策略以中药Senkyu的质量评价为应用场景,成功鉴定出3种新型代谢物(丁烷-1,2,3-三醇、3-脱氧葡糖醛酮和帕拉金糖醇),并通过对照标准品完成了验证。所有工具及经过整理的公开EI质谱数据库均可在RIKEN PRIMe官网(http://prime.psc.riken.jp)的“基于质谱的计算代谢组学”板块免费获取。



