Dataset and trained models for Gated Modulation Protein Language Model for Prediction of Antimicrobial Peptide Activity
收藏资源简介:
This dataset supplements our research work entitled “Gated Protein Language Modeling for Accurate Prediction of Antimicrobial Peptide Activity.” The database follows the directory structure described below: five_fold_ecoli:Computational datasets constructed for model training, stored in .pkl format. five_fold_s_aureus:Five-fold cross-validation datasets for Staphylococcus aureus. five_fold_p_aeruginosa: Five-fold cross-validation datasets for P. Aeruginosa embed:Residue-level peptide embeddings for E. coli and S. aureus and P. aeruginosa peptides computed using Protein Language models.. grampa.csv:The original peptide sequence database containing antimicrobial peptides across multiple bacterial species, including E. coli, P. Aeruginosa, and S. aureus. All derived datasets were constructed from this base file. The original grampa.csv dataset is attributed to the study: Deep learning regression model for antimicrobial peptide designhttps://www.biorxiv.org/content/10.1101/692681v1.full



