遇见数据集

FrustraMPNN: Full training runs/weights + Complete single-residue frustration data (predicted/calculated) for E. coli/Human AF validation datasets

收藏
Zenodo2026-01-24 更新2026-05-26 收录
官方服务:

资源简介:

Introduction This Zenodo entry contains all the generated and analysed data for the Paper "FrustraMPNN: An ultra-fast deep learning tool for proteome-scale analysis of deep mutational single residue local energetic frustration profiles in proteins" (DOI preprint: https://doi.org/10.64898/2026.01.22.701012). In addition to the data needed to reproduce the results from the paper, this entry also contains full single-residue frustration data for approximately 4,300 E. coli protein structures and approximately 3,400 selected proteins from the human proteome. Both validation sets contain roughly 25 Mio datapoints for single-residue frustration indices, and frustration indices are calculated using the wrapped version of FrustratometeR (see frustrapy) and FrustraMPNN. Contents The text files afvalid_human_model_ids.txt and aftest_ecoli_model_ids.txt contains all ids for the tested alphafold structures used in this study for different purposes (see the paper for more details). The PDBs were not altered in any way and can be downloaded from the AlphaFold database. The following table lists all zip files added to this Zenodo entry with a short description about their content. File name (zip) Description full_training_runs_and_weights Contains all training logs and weights produced in this study, split by the dataset used for training (FireProt, MegaScale). The weights and logs are named by the different settings used with the following naming scheme, and possible values: [DATASET]_[DATAPOINTS]_[ADD_LT]_[ADD_ESM]_[SPLIT]_epoch=[XX]_val_frustration_spearman=[VAL].ckpt DATASET: fireprot, megascale DATAPOINTS: default, all_muts, equal balancing strategy during training (see paper for details) ADD_LT: la-true, la-false Light attention added during training (bool) ADD_ESM: esm-True, esm-False ESM embeddings added during training (bool) SPLIT: 1 - 5 Split number for 5-fold cross-validation XX: 1-100 epoch number for early stopping VAL: 0.00 - 1.00 Spearman correlation for the validation set, only the best model for each run is saved training_data Contains all the data used to train the neural networks, including the raw ddG CSVs from the ThermoMPNN paper and processed Fireprot and MegaScale datasets (including PDBs). Both folders "fireprot" and "megascale" contain: training, validation, and test splits for different datapoint variants (default set, saturated set, equally balanced set) CSVs for the default (called "full_data"), all possible single-point mutations (called "all_mutations"), and the equally balanced sets (called "all_mutations_[...]_equal_category") single-residue frustration data (pkl files, additional zip folder) for all PDBs, which are used in the respective datasets jupyter notebooks for data processing steps afvalidation Contains around 3,400 entries for the external Alphafold validation set used in the paper, randomly selected from the human genome; 50-1,000 amino acids long. The zip file contains Complete calculated single-residue frustration data for all proteins (pkls_single-res_21cpu) using FrustratometeR (details see paper) in original pkl format and csv files pkl files --> frustrapy (frustrapy) needs to be installed to correctly parse data Complete prediction of single-residue frustration values using the best weights from the FireProt and MegaScale training runs csv files, for every position and every mutation, the single-residue frustration value is reported Important: 29 pkl files are damaged due to parsing errors during the calculation, but the errors could not be resolved. Therefore, it is recommended to use the csv files of the FrustratometeR calculations. timing_tests Contains the measured inference time using FrustraMPNN and calculation time using 1/10/21 CPUs for the same calculation to measure the speedup factor for the Human Alphafold validation set. Also contains a combined dataframe (ecoli_proteome_timing_approx.csv) for all E. coli proteins examined with approximated times based on a linear fit from the human protein calculations. [10,21,50]cpu_timing.csv: Measured time for (successful) calculated single-residue frustration values using FrustraMPNN (inference column) and X-CPUs for frustrapy (frustration column). timing-frustration-binned_ALL.csv: combined dataframe for the csv above with additional times for 1CPU (column label) calculations ecoli_proteome_timing_approx.csv: Approximated times for around 4,300 (proteome-like) E. coli proteins based on a linear fit from the measured times from timing-frustration-binned_ALL.csv enzyms_example Contains input and output data for the three selected enzymes shown in the paper as example cases (PDB ids: 1CBG, 1NLU, and 4BLM). Each output folder contains: cleaned pdb structure raw single residue frustration data (pkl files) for all positions and all mutations of the processed structure inference results as csv combined results (predicted and calculated single-residue frustration data combined) location_secstruct R^2 and Spearman correlation values for the analysis of location (surface, core, boundary) and secondary structure (helix, sheet, loops) regarding the prediction of frustration categories. per_aa_error RMSE, Spearman correlation, Accuracy, and F1-score for all possible individual mutations. Global metrics are reported, along with values for the three frustration categories (neutral, minimal, and highly frustrated) based on the external AlphaFold validation benchmark set from the human proteome. ecoli_test_set The same structure as the folder for afvalidation, but with an extra folder "example_scatter" as the underlying data for Figure S4 in the corresponding paper to this zenodo entry. Also contains around 4,300 predicted and calculated single-residue frustration datapoints.

提供机构:
Zenodo
创建时间:
2026-01-24
二维码
社区交流群
二维码
科研交流群
商业服务