CropStabDB: data, trained models and validation for a crop protein stability resource across 30 plant species
收藏资源简介:
CropStabDB is a genome-scale resource for plant protein stability and structure-guided protein engineering, covering 1,425,716 proteins across 30 plant species (29 crops and Arabidopsis thaliana, 12 families), each scored by the CropStab Score and decomposed into six interpretable components. This archive contains the complete data, trained models and validation underlying the resource:- master feature table (1,425,716 proteins x 149 sequence-derived features) with CropStab Scores;- the Engineering Atlas of 2,362,229 candidate stabilising substitutions (541,987 high-confidence at ECS >= 0.7) and per-protein summaries;- trained Random Forest and XGBoost models with predictions and SHAP values;- the SQLite database and parquet table behind the cropstab.org web interface;- the 18 validation-layer outputs (AlphaFold pLDDT, N-end rule, ortholog conservation, published families, measured Kd, and more);- manuscript figures and supplementary tables. Source code: https://github.com/Israfil-Hossen/cropstabdbWeb interface: https://cropstab.org
CropStabDB是一款面向植物蛋白质稳定性与结构导向蛋白质工程的基因组级资源,涵盖30个植物物种(29种作物及拟南芥(Arabidopsis thaliana),隶属于12个科)的1,425,716条蛋白质序列,每条序列均通过CropStab Score进行评分,并被拆解为6个可解释的组分。 该存档包含支撑该资源的完整数据集、训练模型与验证集: - 主特征表(1,425,716条蛋白质 × 149个序列衍生特征),附带CropStab Score评分; - 包含2,362,229个候选稳定性增强替换突变的工程图谱(其中541,987个在ECS≥0.7时为高置信度结果),以及单蛋白质汇总信息; - 带有预测结果与SHAP值的训练随机森林(Random Forest)与XGBoost模型; - 支撑cropstab.org网页界面的SQLite数据库与Parquet数据表; - 18个验证层输出结果,包括AlphaFold pLDDT、N端规则、同源保守性、已发表蛋白家族、实测Kd值等; - 论文配图与补充表格。 源代码:https://github.com/Israfil-Hossen/cropstabdb 网页界面:https://cropstab.org



