Accurate protein stability prediction for small domains using mega-scale experiments
收藏资源简介:
Accurate protein stability prediction for small domains using mega-scale experiments IMPORTANT! Please register your use of these data so that we can continue to release new useful datasets! ==========================================================================Links========================================================================== Source code: https://github.com/yehlincho/absolute-stability-predictorModel weights: https://huggingface.co/Yehlin/absolute-stabilityData access: https://forms.gle/4ZnXZSnTBvaykkAi9Training structures dataset (train/val/test) in mgnify_training_index.csv: https://drive.google.com/drive/folders/1SPHsaC2tcwtnQuU0gbZeEFUuol-bhSZM ========================================================================== Each CSV contains identifiers (name/pdb/ted_id), experimental values,and one predicted_<model> column per model evaluated. ==========================================================================TOP-LEVEL FILES (Training / Raw Datasets)========================================================================== 230515_K50dG_dmsv4_dmsv5_dmsv7_concat260429.csv Full MGnify Stability dataset with recalibrated dG values (dmsv4 + dmsv5 + dmsv7) 220122_dmsv4_count_table.csv DMSv4 raw NGS count table 220331_dmsv5_count_table.csv DMSv5 raw NGS count table mgnify_training_index.csv MGnify sequence index used for training dmsv4_filtered_train_splits.csv DMSv4 training split after unstructured terminus filtering ==========================================================================benchmarks.zip========================================================================== S1724_PDB_sequences.csv S1724 benchmark: PDB sequences for 35 wild-type proteins S1724_thermomutdb_cleaned_withseq.csv S1724 benchmark: curated ThermoMutDB stability measurements with sequences TED_under100aa_filtered_20percent_5k.csv TED domains <100 aa, filtered to 20% sequence identity, 5k cap per temperature bin TED_100to200aa_filtered_1point3mill.csv TED domains 100-200 aa, ~1.3M sequences ==========================================================================dmsv7_analysis.zip========================================================================== Complete DMSv7 analysis folder (notebooks, scripts, intermediate results) ==========================================================================scripts_to_calculate_K50_and_dG.zip========================================================================== Collection of Jupyter notebooks and analysis scripts used to process unfolded datasets and calculate K50 and ΔG values for the libraries. ==========================================================================figures.zip -- per-figure prediction CSVs========================================================================== Figure 1 & 2 -- Mgnify stability benchmarkfig1_2__mgnify_stability_predictions.csv Mgnify test set. Columns: name, seq, experiment (dG), predicted_Single_SaProtdG, predicted_Single_ESM3dG, predicted_BioEMU, predicted_IFUM, predicted_ESM3dG, predicted_Aug_ESM3dG, predicted_SaProtdG, predicted_Aug_SaProtdG Figure 3 -- Megascale point mutantsfig3__megascale_point_mutants_predictions.csv Megascale ddG test set. Columns: name, experiment (ddG), predicted_ESM3dG, predicted_SaProtdG, predicted_Single_ESM3dG, predicted_Single_SaProtdG, predicted_ProteinDPO, predicted_ThermoMPNN, aa_seq, wt_seq Figure 3 -- Mgnify insertion/deletion benchmarkfig3__mgnify_insdel_predictions.csv MGnify insdel test set. Columns: name, wt_name, experiment (ddG), predicted_ESM3dG, predicted_SaProtdG, predicted_Single_ESM3dG, predicted_Single_SaProtdG, predicted_ProteinDPO, aa_seq Figure 3 -- Megascale insertion/deletion benchmarkfig3__megascale_insdel_predictions.csv Megascale insdel test set. Columns: pdb, wt_name, experiment (ddG), aa_seq_full, aa_seq, predicted_ESM3dG, predicted_SaProtdG, predicted_Rosetta Figure 4 -- ThermoMut stability benchmarkfig4__thermomut_predictions.csv ThermoMut (S1724) test set. Columns: name, experiment (ddG), experiment_dG, dG_wt, mut_aa, wt_aa, position, paper_seq, GJR_trim_seq, PDB, predicted_SaProtdG, predicted_ddG_SaProtdG, predicted_SaProtdG_trim, predicted_ddG_SaProtdG_trim, predicted_Aug_SaProtdG, predicted_ddG_Aug_SaProtdG, predicted_Aug_SaProtdG_trim, predicted_ddG_Aug_SaProtdG_trim, predicted_ESM3dG, predicted_ddG_ESM3dG, predicted_ESM3dG_trim, predicted_ddG_ESM3dG_trim, predicted_Aug_ESM3dG, predicted_ddG_Aug_ESM3dG, predicted_Aug_ESM3dG_trim, predicted_ddG_Aug_ESM3dG_trim, predicted_SaProtdG_no_sigmoid, predicted_ddG_SaProtdG_no_sigmoid, predicted_ThermoMPNN Figure 5 -- TED thermophilicity gradientfig5__TED_temperature_predictions.csv TED domains with organism growth temperature. Columns: ted_id, UniProtID, temperature, cath_label, organism, predicted_Augmented_ESM3dG, predicted_Augmented_ESM3dG_ddG, predicted_ESM_IF_ddG, predicted_ProteinMPNN_CE_scaled, predicted_AF2_pLDDT_scaled, aa_seq Figure 6 -- De novo designs: DMSV2 (Cho et al. 2025)fig6__dmsv2_denovo_predictions.csv De novo designs from Cho et al. (2025). Columns: pdb, experiment (dG), predicted_Augmented_ESM3dG, aa_seq Figure 6 -- Nanobody thermal stabilityfig6__nanobody_predictions.csv Nanobody Tm dataset. Columns: name, experiment_Tm (C), predicted_ESM3dG, predicted_Boltz2_pLDDT, aa_seq Figure 6 -- Rosetta & Dark-Matter de novo designsfig6__rosetta_dark_matter_designs_predictions.csv Rosetta (2012) and Dark-Matter fold designs. Columns: pdb, dataset, experiment (Success/Failure), predicted_Augmented_ESM3dG, predicted_ESM3dG, predicted_AF2_pLDDT, predicted_ProteinMPNN_NegCE, aa_seq Figure 6 -- RFdiffusion binder designsfig6__rfdiffusion_binders_predictions.csv RFdiffusion binder designs across 7 targets. Columns: pdb, target, experiment (binder label), predicted_Augmented_ESM3dG, predicted_ESM3dG, predicted_ProteinMPNN



