遇见数据集

Dataset for Peptide Activity Prediction (Minimum Inhibitory Concentration Prediction) -

收藏
Zenodo2026-02-13 更新2026-05-26 收录
官方服务:

资源简介:

This dataset supplements our research work entitled “Gated Protein Language Modeling for Accurate Prediction of Antimicrobial Peptide Activity.” The database follows the directory structure described below: alphafold_pdb_ecoli:PDB structures generated using AlphaFold2 for antimicrobial peptides targeting Escherichia coli. alphafold_pdb_stap:PDB structures generated using AlphaFold2 for antimicrobial peptides targeting Staphylococcus aureus. fasta_ecoli, fasta_stap:FASTA-formatted peptide sequences for E. coli and S. aureus. five_fold_ecoli:Computational datasets constructed for model training, stored in .pkl format.Each .pkl file contains five cross-validation splits. Multiple .pkl files are provided with slightly different feature configurations; please refer to readme.txt for details.ecoli_dataset.pkl is the final dataset used for ModProt model training. five_fold_s_aureus:Five-fold cross-validation datasets for Staphylococcus aureus. protT5:Residue-level peptide embeddings for E. coli and S. aureus peptides computed using ProtT5. grampa.csv:The original peptide sequence database containing antimicrobial peptides across multiple bacterial species, including E. coli and S. aureus. All derived datasets were constructed from this base file. The original grampa.csv dataset is attributed to the study: Deep learning regression model for antimicrobial peptide designhttps://www.biorxiv.org/content/10.1101/692681v1.full

本数据集补充了题为《Gated Protein Language Modeling for Accurate Prediction of Antimicrobial Peptide Activity》(门控蛋白质语言模型用于抗菌肽活性精准预测)的研究工作。 该数据库遵循如下目录结构: alphafold_pdb_ecoli:收录针对大肠杆菌(Escherichia coli)的抗菌肽(Antimicrobial Peptide)的PDB结构文件,均由AlphaFold2生成。 alphafold_pdb_stap:收录针对金黄色葡萄球菌(Staphylococcus aureus)的抗菌肽的PDB结构文件,均由AlphaFold2生成。 fasta_ecoli、fasta_stap:收录大肠杆菌与金黄色葡萄球菌的肽序列,格式为FASTA。 five_fold_ecoli:收录用于模型训练的计算数据集,以.pkl格式存储。每个.pkl文件包含5折交叉验证(cross-validation)划分;本次提供了多个特征配置略有差异的.pkl文件,详细信息请参阅readme.txt。其中ecoli_dataset.pkl为用于ModProt模型训练的最终数据集。 five_fold_s_aureus:收录针对金黄色葡萄球菌的5折交叉验证数据集。 protT5:收录使用ProtT5计算得到的、针对大肠杆菌与金黄色葡萄球菌肽序列的残基级肽嵌入向量。 grampa.csv:原始肽序列数据库,收录涵盖多种细菌物种的抗菌肽,包含大肠杆菌与金黄色葡萄球菌;所有衍生数据集均基于该基础文件构建。 原始grampa.csv数据集源自以下研究: 《用于抗菌肽设计的深度学习回归模型》(Deep learning regression model for antimicrobial peptide design),链接:https://www.biorxiv.org/content/10.1101/692681v1.full

提供机构:
Zenodo
创建时间:
2026-02-13
二维码
社区交流群
二维码
科研交流群
商业服务