Rate-Limiting Steps and Machine-Learning Prediction of Enzyme Turnover Numbers: A Benchmark Reanalysis Across 21 Enzyme Families" |
收藏资源简介:
This archive contains the data, code, and intermediate results for the manuscript "Rate-Limiting Steps and Machine-Learning Prediction of Enzyme Turnover Numbers: A Benchmark Reanalysis Across 21 Enzyme Families" (Aktaş & Rahayu, submitted to the Journal of Biomolecular Structure and Dynamics). What the study does. Machine-learning models of enzyme turnover number (kcat) are reported to generalize poorly to enzymes unlike their training data. We ask whether prediction accuracy differs between enzymes whose rate-limiting step is chemical and enzymes whose rate-limiting step is non-chemical (for example, conformational change or product release), and whether any difference depends on model architecture. We reanalyzed the published predictions of five models (DLKcat, UniKP, EITLEM-Kinetics, CataPro, CatPred) for 22 enzyme families from the public EnzyArena benchmark. Twenty-one families were classified from independent literature evidence (16 chemistry-limited, 5 dynamics-limited); one was retained as uncertain and excluded from all tests. No model was trained or modified. Main results, as reported in the manuscript. In a precision-weighted, model-pooled analysis, chemistry-limited families showed higher Spearman correlations than dynamics-limited families for the two sequence-only models (p = 0.0012), but not for the two models with structural-type input (p = 0.63). The independence-respecting family-level comparison, which we treat as the primary basis for judging significance, was directionally consistent but not significant (Mann-Whitney p = 0.354; Cohen's d = 0.78). The model-pooled result weakens to non-significance when only the strongest mechanistic evidence is used or when the two largest dynamics-limited families are excluded. Model architecture, training data, and representation are confounded, so the results are consistent with, but do not establish, the proposed explanation. Contents.- `data/`: the three EnzyArena source files (unmodified snapshot, September 2026) and the original extraction script.- `scripts/`: `extract_all_families.py` (recomputes per-family Spearman rho for the 20 EnzyArena-derived families from the raw CSVs), `recompute_ak_ppo.py` (adenylate kinase and polyphenol oxidase), `sensitivity_analyses.py` (manuscript Tables 2, 4 and 5, pLM-axis regrouping, mixed-effects models), `compute_r2_full.py` (R² and RMSE), and figure scripts.- `final_audit_recomputations.py`: family-level tests, bootstrap, and leave-one-family-out analysis.- `results/`: per-family correlation files and R²/RMSE results. `results/legacy_17family/` holds superseded outputs from an earlier 17-family version of the analysis, kept for provenance only.- `figures/final/`: the six manuscript figures. `figures/legacy_17family/` holds superseded figures.- `README.md`: directory guide, run instructions, and a table linking each manuscript item to its source file. Data and licensing note. The EnzyArena repository (github.com/zishuozeng/EnzyArena) is MIT-licensed; the MIT license here covers the authors' code. We make no claim about the licensing of the underlying BRENDA- and SABIO-RK-derived data; please consult those databases' own terms. Known limitations.(1) For five families, EnzyArena holds several kcat campaigns; rho was averaged across campaigns and n is the total number of mutants, which may overstate precision (a sensitivity check is included). (2) For thymidylate synthase, EnzyArena contains no kcat row for the cited reference, so a kcat/Km row was used (a sensitivity check excluding it is included). (3) The mixed-effects model with crossed family and model random effects reports a non-positive-definite Hessian, so its standard error is approximate. (4) Figure scripts write to the authors' original working directory and need their output path edited before running.



