遇见数据集

Hydroxylase Thermostability Prediction Based on Self-Trained Semisupervised Iteration and Bayesian Dynamic Tuning

收藏
NIAID Data Ecosystem2026-05-10 收录
官方服务:

资源简介:

Current enzyme thermostability prediction models are predominantly designed for cross-family generalization, with limited focus on hydroxylases, which restricts their accuracy and applicability in hydroxylase-specific thermostability design. In this study, we develop HyS-BST, a dedicated self-trained semisupervised framework for hydroxylase thermostability prediction. Leveraging a limited hydroxylase data set, HyS-BST integrates a self-training strategy with Bayesian dynamic tuning to achieve high-precision prediction of mutant thermostability in terms of ΔΔG. Experimental results demonstrate that after only ten training iterations, HyS-BST attains a coefficient of determination (R2) of 0.96, a Pearson correlation coefficient (PCC) of 0.98, and a root mean squared error (RMSE) as low as 0.06 on the test set. Compared with the optimal cross-family generalization model, HyS-BST improves PCC and RMSE by approximately 70%. Overall, this framework provides a specialized, efficient, and cost-effective solution for hydroxylase thermostability prediction, substantially reducing the candidate search space and experimental resources required for downstream validation.

当前的酶热稳定性预测模型大多针对跨家族泛化性设计,对羟化酶(hydroxylases)的关注较为有限,这限制了模型在羟化酶专属热稳定性设计中的准确性与应用潜力。本研究开发了HyS-BST——一种专为羟化酶热稳定性预测设计的自训练半监督框架。依托有限的羟化酶数据集,HyS-BST将自训练策略与贝叶斯动态调优相结合,实现了以ΔΔG为评价指标的突变体热稳定性高精度预测。实验结果表明,仅经过10次训练迭代后,HyS-BST在测试集上即可取得0.96的决定系数(coefficient of determination,R²)、0.98的皮尔逊相关系数(Pearson correlation coefficient,PCC),以及低至0.06的均方根误差(root mean squared error,RMSE)。与最优的跨家族泛化模型相比,HyS-BST的PCC与RMSE性能分别提升约70%。总体而言,该框架为羟化酶热稳定性预测提供了专属、高效且具成本效益的解决方案,可大幅缩减下游验证所需的候选搜索空间与实验资源。

创建时间:
2026-03-05
二维码
社区交流群
二维码
科研交流群
商业服务