EnzyBase12k
收藏资源简介:
EnzyBase12k: A curated enzyme dataset for catalytic pH optimum prediction EnzyBase12k is a curated dataset integrating structural and sequential information with experimentally determined catalytic pH optima. It contains 11,615 enyzmes structures (CIF format) predicted with AlphaFold with or without template mode. The EnzyBase12k_metadata.csv contains UniProt ID, (cut/uncut) UniProt Sequence, catalytic pH optimum, structure source, if PDB structure matching to the UniProt sequence was available and further information. The dataset contains monomers and homopolymers; heteromers were excluded. Homopolymers were reduced to one enzyme subunit and handledcas monomer. ### License This dataset integrates third-party data from the following sources: UniProt sequences were obtained from https://www.uniprot.org/ [CC BY 4.0] PDB structures were obtained from https://rcsb.org [CC0 1.0 Universal] AlphaFold2 models were downloaded from https://alphafold.ebi.ac.uk/ [CC BY 4.0] AlphaFold3 models predicted using AlphaFold v3.0.1, https://github.com/google-deepmind/alphafold3 [CC-BY-NC-SA 4.0] The metadata ('EnzyBase12k_metadata.csv') and the dataset curation process (including filtering, structure selection and sequence mapping)are licensed by the author under [CC BY-NC 4.0]. --- *Note: This dataset was curated as part of the Master's thesis 'Development of Graph Neural Networks (GNNs) for Predicting Enzymatic pH Optima' and is not yet peer-reviewed.*



