UniPepBert
收藏资源简介:
This Zenodo record contains the datasets curated, processed, and used in the study “PepUniBERT: Length-Stratified Pretraining and Dataset-Scale-Aware Augmentation for Peptide Prediction”. The deposited files include raw data, preprocessed data, and fine-tuning datasets used for peptide function prediction experiments, including pretraining data derived from UniRef50 and downstream peptide bioactivity datasets used for binary classification benchmarks. These files support the pretraining, fine-tuning, data augmentation, benchmark evaluation, and independent validation analyses reported in the manuscript. Source code and environment configuration files are not included in this Zenodo record and are openly available in the accompanying Gitee repository: https://gitee.com/zhao-chengcaas/PepUniBert. Details of data sources, preprocessing procedures, dataset splits, and environment setup are provided in the manuscript, repository documentation, and this Zenodo record. The CC BY 4.0 license applies to author-generated dataset organization, preprocessing outputs, split files, and documentation where applicable. Third-party datasets remain subject to the copyright, license terms, and citation requirements of their original providers. Users should cite the original data sources and comply with the corresponding license terms when reusing or redistributing third-party datasets.



