Prediction of S-Glutathionylation Sites Based on Protein Sequences
收藏资源简介:
S-glutathionylation, the reversible formation of mixed disulfides between glutathione(GSH) and cysteine residues in proteins, is a specific form of post-translational modification that plays important roles in various biological processes, including signal transduction, redox homeostasis, and metabolism inside cells. Experimentally identifying S-glutathionylation sites is labor-intensive and time consuming, whereas bioinformatics methods provide an alternative way to this problem by predicting S-glutathionylation sites in silico. The bioinformatics approaches give not only candidate sites for further experimental verification but also bio-chemical insights into the mechanism of S-glutathionylation. In this paper, we firstly collect experimentally determined S-glutathionylated proteins and their corresponding modification sites from the literature, and then propose a new method for predicting S-glutathionylation sites by employing machine learning methods based on protein sequence data. Promising results are obtained by our method with an AUC (area under ROC curve) score of 0.879 in 5-fold cross-validation, which demonstrates the predictive power of our proposed method. The datasets used in this work are available at http://csb.shu.edu.cn/SGDB.
S-谷胱甘肽化(S-glutathionylation)是谷胱甘肽(glutathione, GSH)与蛋白质半胱氨酸残基之间可逆形成混合二硫键的过程,属于一类特殊的翻译后修饰形式,在细胞内信号转导、氧化还原稳态及代谢等多种生物学过程中发挥关键作用。实验鉴定S-谷胱甘肽化位点往往耗费大量人力与时间,而生物信息学方法可通过虚拟(in silico)预测S-谷胱甘肽化位点,为该问题提供了高效的替代解决方案。此类生物信息学手段不仅可为后续实验验证提供候选修饰位点,还能为解析S-谷胱甘肽化的分子机制提供生化层面的研究视角。本文首先从已发表文献中搜集经实验证实的S-谷胱甘肽化蛋白质及其对应的修饰位点,随后基于蛋白质序列数据,采用机器学习方法构建了一种全新的S-谷胱甘肽化位点预测模型。在5折交叉验证中,本方法取得了0.879的ROC曲线下面积(AUC, area under ROC curve)得分,预测结果表现优异,充分证明了所提模型的预测效能。本研究使用的数据集可从http://csb.shu.edu.cn/SGDB获取。



