Pretraining data for PeptideCLM (UPDATED)
收藏官方服务:
资源简介:
This version update includes changes to Generated_peptides.csv to fix cyclization. The prior upload did not have ring closures generated correctly as SMILES strings. The model in the publication was trained on the dataset containing errors, however to support the community we decided it would be best to release a 10M peptide SMILES dataset for use in future pretraining applications. All strings should now load correctly to mol files with RDKit.
本次版本更新对Generated_peptides.csv文件进行了修改,以修复环化相关问题。此前上传的版本中,环闭合结构未能以正确的SMILES(Simplified Molecular Input Line Entry System)字符串形式生成。本研究发表时所用的模型是在存在错误的数据集上训练得到的,但为支持社区发展,我们决定发布一个包含1000万条肽类SMILES字符串的数据集,以供未来的预训练应用使用。目前所有字符串均可通过RDKit正确加载为mol文件。
提供机构:
Zenodo创建时间:
2024-11-20



