遇见数据集

Accurate protein stability prediction for small domains using mega-scale experiments

收藏
Zenodo2026-06-04 更新2026-05-26 收录
官方服务:

资源简介:

Accurate protein stability prediction for small domains using mega-scale experiments IMPORTANT! Please register your use of these data so that we can continue to release new useful datasets! ==========================================================================Links========================================================================== Source code: https://github.com/yehlincho/absolute-stability-predictorModel weights: https://huggingface.co/Yehlin/absolute-stabilityData access: https://forms.gle/4ZnXZSnTBvaykkAi9Training structures dataset (train/val/test) in mgnify_training_index.csv: https://drive.google.com/drive/folders/1SPHsaC2tcwtnQuU0gbZeEFUuol-bhSZM ========================================================================== Each CSV contains identifiers (name/pdb/ted_id), experimental values,and one predicted_<model> column per model evaluated. ==========================================================================TOP-LEVEL FILES (Training / Raw Datasets)========================================================================== 230515_K50dG_dmsv4_dmsv5_dmsv7_concat260429.csv Full MGnify Stability dataset with recalibrated dG values (dmsv4 + dmsv5 + dmsv7) 220122_dmsv4_count_table.csv DMSv4 raw NGS count table 220331_dmsv5_count_table.csv DMSv5 raw NGS count table mgnify_training_index.csv MGnify sequence index used for training dmsv4_filtered_train_splits.csv DMSv4 training split after unstructured terminus filtering ==========================================================================benchmarks.zip========================================================================== S1724_PDB_sequences.csv S1724 benchmark: PDB sequences for 35 wild-type proteins S1724_thermomutdb_cleaned_withseq.csv S1724 benchmark: curated ThermoMutDB stability measurements with sequences TED_under100aa_filtered_20percent_5k.csv TED domains <100 aa, filtered to 20% sequence identity, 5k cap per temperature bin TED_100to200aa_filtered_1point3mill.csv TED domains 100-200 aa, ~1.3M sequences ==========================================================================dmsv7_analysis.zip========================================================================== Complete DMSv7 analysis folder (notebooks, scripts, intermediate results) ==========================================================================scripts_to_calculate_K50_and_dG.zip========================================================================== Collection of Jupyter notebooks and analysis scripts used to process unfolded datasets and calculate K50 and ΔG values for the libraries. ==========================================================================figures.zip -- per-figure prediction CSVs========================================================================== Figure 1 & 2 -- Mgnify stability benchmarkfig1_2__mgnify_stability_predictions.csv Mgnify test set. Columns: name, seq, experiment (dG), predicted_Single_SaProtdG, predicted_Single_ESM3dG, predicted_BioEMU, predicted_IFUM, predicted_ESM3dG, predicted_Aug_ESM3dG, predicted_SaProtdG, predicted_Aug_SaProtdG Figure 3 -- Megascale point mutantsfig3__megascale_point_mutants_predictions.csv Megascale ddG test set. Columns: name, experiment (ddG), predicted_ESM3dG, predicted_SaProtdG, predicted_Single_ESM3dG, predicted_Single_SaProtdG, predicted_ProteinDPO, predicted_ThermoMPNN, aa_seq, wt_seq Figure 3 -- Mgnify insertion/deletion benchmarkfig3__mgnify_insdel_predictions.csv MGnify insdel test set. Columns: name, wt_name, experiment (ddG), predicted_ESM3dG, predicted_SaProtdG, predicted_Single_ESM3dG, predicted_Single_SaProtdG, predicted_ProteinDPO, aa_seq Figure 3 -- Megascale insertion/deletion benchmarkfig3__megascale_insdel_predictions.csv Megascale insdel test set. Columns: pdb, wt_name, experiment (ddG), aa_seq_full, aa_seq, predicted_ESM3dG, predicted_SaProtdG, predicted_Rosetta Figure 4 -- ThermoMut stability benchmarkfig4__thermomut_predictions.csv ThermoMut (S1724) test set. Columns: name, experiment (ddG), experiment_dG, dG_wt, mut_aa, wt_aa, position, paper_seq, GJR_trim_seq, PDB, predicted_SaProtdG, predicted_ddG_SaProtdG, predicted_SaProtdG_trim, predicted_ddG_SaProtdG_trim, predicted_Aug_SaProtdG, predicted_ddG_Aug_SaProtdG, predicted_Aug_SaProtdG_trim, predicted_ddG_Aug_SaProtdG_trim, predicted_ESM3dG, predicted_ddG_ESM3dG, predicted_ESM3dG_trim, predicted_ddG_ESM3dG_trim, predicted_Aug_ESM3dG, predicted_ddG_Aug_ESM3dG, predicted_Aug_ESM3dG_trim, predicted_ddG_Aug_ESM3dG_trim, predicted_SaProtdG_no_sigmoid, predicted_ddG_SaProtdG_no_sigmoid, predicted_ThermoMPNN Figure 5 -- TED thermophilicity gradientfig5__TED_temperature_predictions.csv TED domains with organism growth temperature. Columns: ted_id, UniProtID, temperature, cath_label, organism, predicted_Augmented_ESM3dG, predicted_Augmented_ESM3dG_ddG, predicted_ESM_IF_ddG, predicted_ProteinMPNN_CE_scaled, predicted_AF2_pLDDT_scaled, aa_seq Figure 6 -- De novo designs: DMSV2 (Cho et al. 2025)fig6__dmsv2_denovo_predictions.csv De novo designs from Cho et al. (2025). Columns: pdb, experiment (dG), predicted_Augmented_ESM3dG, aa_seq Figure 6 -- Nanobody thermal stabilityfig6__nanobody_predictions.csv Nanobody Tm dataset. Columns: name, experiment_Tm (C), predicted_ESM3dG, predicted_Boltz2_pLDDT, aa_seq Figure 6 -- Rosetta & Dark-Matter de novo designsfig6__rosetta_dark_matter_designs_predictions.csv Rosetta (2012) and Dark-Matter fold designs. Columns: pdb, dataset, experiment (Success/Failure), predicted_Augmented_ESM3dG, predicted_ESM3dG, predicted_AF2_pLDDT, predicted_ProteinMPNN_NegCE, aa_seq Figure 6 -- RFdiffusion binder designsfig6__rfdiffusion_binders_predictions.csv RFdiffusion binder designs across 7 targets. Columns: pdb, target, experiment (binder label), predicted_Augmented_ESM3dG, predicted_ESM3dG, predicted_ProteinMPNN

基于大规模实验的小型结构域蛋白质稳定性精准预测 重要提示!请登记您对本数据集的使用情况,以便我们持续发布更多实用数据集! ==========================================================================链接========================================================================== 源代码:https://github.com/yehlincho/absolute-stability-predictor 模型权重:https://huggingface.co/Yehlin/absolute-stability 数据获取:https://forms.gle/4ZnXZSnTBvaykkAi9 ========================================================================== 每个CSV文件均包含标识符(名称/PDB(Protein Data Bank)/ted_id)、实验数值,以及对应每个评估模型的predicted_<model>列。 ==========================================================================顶层文件(训练/原始数据集)========================================================================== 230515_K50dG_dmsv4_dmsv5_dmsv7_concat260429.csv:完整MGnify稳定性数据集,包含重新校准后的dG值(dmsv4 + dmsv5 + dmsv7) 220122_dmsv4_count_table.csv:DMSv4原始NGS(Next-Generation Sequencing,下一代测序)计数表 220331_dmsv5_count_table.csv:DMSv5原始NGS计数表 mgnify_training_index.csv:训练所用的MGnify序列索引文件 dmsv4_filtered_train_splits.csv:经过非结构化末端过滤后的DMSv4训练划分集 ==========================================================================benchmarks.zip========================================================================== S1724_PDB_sequences.csv:S1724基准测试集:35种野生型蛋白质的PDB序列 S1724_thermomutdb_cleaned_withseq.csv:S1724基准测试集:经整理的ThermoMutDB稳定性测量数据及对应序列 TED_under100aa_filtered_20percent_5k.csv:氨基酸长度小于100的TED结构域,经序列同一性过滤至20%,每个温度分组的样本量上限为5000 TED_100to200aa_filtered_1point3mill.csv:氨基酸长度介于100到200之间的TED结构域,约含130万条序列 ==========================================================================dmsv7_analysis.zip========================================================================== 完整的DMSv7分析文件夹(包含Jupyter笔记本、脚本及中间结果) ==========================================================================figures.zip -- 单图预测结果CSV文件========================================================================== 图1与图2 —— MGnify稳定性基准测试 fig1_2__mgnify_stability_predictions.csv:MGnify测试集,列包括:name, seq, experiment (dG), predicted_Single_SaProtdG, predicted_Single_ESM3dG, predicted_BioEMU, predicted_IFUM, predicted_ESM3dG, predicted_Aug_ESM3dG, predicted_SaProtdG, predicted_Aug_SaProtdG 图3 —— 大规模单点突变测试 fig3__megascale_point_mutants_predictions.csv:大规模ddG测试集,列包括:name, experiment (ddG), predicted_ESM3dG, predicted_SaProtdG, predicted_Single_ESM3dG, predicted_Single_SaProtdG, predicted_ProteinDPO, predicted_ThermoMPNN, aa_seq, wt_seq 图3 —— MGnify插入/缺失基准测试 fig3__mgnify_insdel_predictions.csv:MGnify插入缺失测试集,列包括:name, wt_name, experiment (ddG), predicted_ESM3dG, predicted_SaProtdG, predicted_Single_ESM3dG, predicted_Single_SaProtdG, predicted_ProteinDPO, aa_seq 图3 —— 大规模插入缺失基准测试 fig3__megascale_insdel_predictions.csv:大规模插入缺失测试集,列包括:pdb, wt_name, experiment (ddG), aa_seq_full, aa_seq, predicted_ESM3dG, predicted_SaProtdG, predicted_Rosetta 图4 —— ThermoMut稳定性基准测试 fig4__thermomut_predictions.csv:ThermoMut(S1724)测试集,列包括:name, experiment (ddG), experiment_dG, dG_wt, mut_aa, wt_aa, position, paper_seq, GJR_trim_seq, PDB, predicted_SaProtdG, predicted_ddG_SaProtdG, predicted_SaProtdG_trim, predicted_ddG_SaProtdG_trim, predicted_Aug_SaProtdG, predicted_ddG_Aug_SaProtdG, predicted_Aug_SaProtdG_trim, predicted_ddG_Aug_SaProtdG_trim, predicted_ESM3dG, predicted_ddG_ESM3dG, predicted_ESM3dG_trim, predicted_ddG_ESM3dG_trim, predicted_Aug_ESM3dG, predicted_ddG_Aug_ESM3dG, predicted_Aug_ESM3dG_trim, predicted_ddG_Aug_ESM3dG_trim, predicted_SaProtdG_no_sigmoid, predicted_ddG_SaProtdG_no_sigmoid, predicted_ThermoMPNN 图5 —— TED嗜热性梯度测试 fig5__TED_temperature_predictions.csv:携带生物体生长温度信息的TED结构域数据集,列包括:ted_id, UniProtID, temperature, cath_label, organism, predicted_Augmented_ESM3dG, predicted_Augmented_ESM3dG_ddG, predicted_ESM_IF_ddG, predicted_ProteinMPNN_CE_scaled, predicted_AF2_pLDDT_scaled, aa_seq 图6 —— 从头设计序列:DMSV2(Cho等人,2025) fig6__dmsv2_denovo_predictions.csv:Cho等人(2025)发布的从头设计序列数据集,列包括:pdb, experiment (dG), predicted_Augmented_ESM3dG, aa_seq 图6 —— 纳米抗体热稳定性测试 fig6__nanobody_predictions.csv:纳米抗体Tm值数据集,列包括:name, experiment_Tm (°C), predicted_ESM3dG, predicted_Boltz2_pLDDT, aa_seq 图6 —— Rosetta与暗物质折叠从头设计序列 fig6__rosetta_dark_matter_designs_predictions.csv:Rosetta(2012)与暗物质折叠设计序列数据集,列包括:pdb, dataset, experiment (Success/Failure), predicted_Augmented_ESM3dG, predicted_ESM3dG, predicted_AF2_pLDDT, predicted_ProteinMPNN_NegCE, aa_seq 图6 —— RFdiffusion结合蛋白设计序列 fig6__rfdiffusion_binders_predictions.csv:覆盖7个靶点的RFdiffusion结合蛋白设计序列数据集,列包括:pdb, target, experiment (binder label), predicted_Augmented_ESM3dG, predicted_ESM3dG, predicted_ProteinMPNN

提供机构:
Zenodo
创建时间:
2026-05-19
二维码
社区交流群
二维码
科研交流群
商业服务