遇见数据集

Horizyn-1 model weights and development dataset: Dual-encoder contrastive learning accelerates enzyme discovery

收藏
Zenodo2026-05-26 更新2026-05-29 收录
官方服务:

资源简介:

Overview This repository provides the official model weights for the Horizyn-1 enzyme discovery models, as well as the open-source development dataset, as described in "Dual-encoder contrastive learning accelerates enzyme discovery." The accompanying code for training, inference, and evaluation is available at: https://github.com/dayhofflabs/horizyn. Model Checkpoints Two model checkpoints are provided in this archive: horizyn_v1_0_inf.ckpt (Inference Model): The primary model recommended for inference. Note: This model was trained on a much larger, withheld dataset and was not trained on the development dataset provided in this repository. horizyn_v1_0_dev.ckpt (Development Model): A model trained entirely on the accompanying open-source development dataset, provided for reproducibility and benchmarking. Development Dataset Overview & Splits The dataset included in this repository was used to train and evaluate the development model (horizyn_v1_0_dev.ckpt). It includes the full set of reaction SMILES, protein-reaction pairs, and pre-computed ProtT5 protein embeddings. The training and test sets were created by splitting on reactions to prevent data leakage. The test set was strictly filtered to exclude any reactions with high similarity to any reactions in the training set (see manuscript for details). For evaluation, the test setup involves identifying the correct enzyme for each test reaction from a total screening pool of 216,132 proteins contained in this dataset. Dataset Statistics Reactions: 10,785 (Train) / 1,012 (Test) Enzymes: 192,769 (Train) / 32,100 (Test) Reaction-Enzyme Pairs: 257,733 (Train) / 33,996 (Test) File Manifest The archive contains the following standardized files: horizyn_v1_0_inf.ckpt & horizyn_v1_0_dev.ckpt Model weights for the Horizyn-1 inference and development models, respectively. train_rxns.csv & test_rxns.csv Contains reaction SMILES strings for the development dataset. Columns: rs_id, reaction_id, reaction_smiles train_pairs.csv & test_pairs.csv Defines the positive training and testing pairs for the development dataset. Columns: pr_id, reaction_id, protein_id prots_t5.h5.gz Compressed HDF5 file containing pre-computed protein embeddings (ProtT5-XL). Structure (when uncompressed): /ids: Dataset of protein IDs (strings) /vectors: Dataset of embeddings (float32, shape: [N, 1024]) prots.fasta.gz (Reference) Compressed FASTA file containing protein sequences corresponding to the embeddings in the HDF5 file. Citation If you use these models or the development dataset, please cite the associated publication: Rocks, J. W., Truong, D. P., Rappoport, D., Maddrell-Mander, S., Martin-Alarcon, D. A., Lee, T. M., Crossan, S., & Goldford, J. E. (2026). Dual-encoder contrastive learning accelerates enzyme discovery. Proc. Natl. Acad. Sci. U.S.A. 123 (12) e2520070123. https://doi.org/10.1073/pnas.2520070123

提供机构:
Zenodo
创建时间:
2026-05-26
二维码
社区交流群
二维码
科研交流群
商业服务