ProtBert and ESM2 trained model and data
收藏资源简介:
This dataset supplements our research work entitled “Gated Protein Language Modeling for Accurate Prediction of Antimicrobial Peptide Activity.” The database follows the directory structure described below: alphafold_pdb_ecoli:PDB structures generated using AlphaFold2 for antimicrobial peptides targeting Escherichia coli. alphafold_pdb_stap:PDB structures generated using AlphaFold2 for antimicrobial peptides targeting Staphylococcus aureus. fasta_ecoli, fasta_stap:FASTA-formatted peptide sequences for E. coli and S. aureus. five_fold_ecoli:Computational datasets constructed for model training, stored in .pkl format.Each .pkl file contains five cross-validation splits. Multiple .pkl files are provided with slightly different feature configurations; please refer to readme.txt for details.ecoli_dataset.pkl is the final dataset used for ModProt model training. five_fold_s_aureus:Five-fold cross-validation datasets for Staphylococcus aureus. protT5:Residue-level peptide embeddings for E. coli and S. aureus peptides computed using ProtT5. grampa.csv:The original peptide sequence database containing antimicrobial peptides across multiple bacterial species, including E. coli and S. aureus. All derived datasets were constructed from this base file. The original grampa.csv dataset is attributed to the study: Deep learning regression model for antimicrobial peptide designhttps://www.biorxiv.org/content/10.1101/692681v1.full



