遇见数据集

Data Archiving and Access for NaFM: Pre-training a Foundation Model for Small-Molecule Natural Products

收藏
DataCite Commons2025-06-01 更新2025-09-08 收录
官方服务:

资源简介:

<b>pretrain_smiles.pkl</b>: Preprocessed data used for model pretraining. The original data was obtained from the COCONUT database: https://coconut.naturalproducts.net/<b>classification_data.csv</b>: Data prepared for the Natural Product Taxonomy Classification experiment. The original dataset was sourced from the following archive: https://zenodo.org/records/5068687#.YOKJQOgzaUl<b>NPClassifier_dataset_refreshed.csv</b>: Data curated for direct comparison with NPClassifier. The original data is available at: https://github.com/mwang87/NP-Classifier/tree/master/training/Data/NPClassifier_dataset.xlsx<b>regression_data.csv</b>: Dataset used for natural product bioactivity prediction tasks. The original data was retrieved from the NPASS database: https://bidd.group/NPASS/<b>lotus_data.csv</b>: Data prepared for biological source prediction and related mining tasks. The source data was collected from the LOTUS database: https://lotus.naturalproducts.net/<b>bgc_data.csv</b>: Dataset constructed for biosynthetic gene cluster mining. The original sources include the MIBiG database (https://mibig.secondarymetabolites.org/) and Pfam (http://pfam.xfam.org/)<b>external_data.csv</b>: Dataset used for bioactivity screening of natural products. The original data was obtained from the NPASS database: https://bidd.group/NPASS/

提供机构:
figshare
创建时间:
2025-05-09
二维码
社区交流群
二维码
科研交流群
商业服务