遇见数据集

Dataset EnzyBase12k: A Curated Dataset for Predicting Enzyme pH Optima Using pHoptNN

收藏
Zenodo2026-03-11 更新2026-05-26 收录
官方服务:

资源简介:

A Curated Dataset for Predicting Enzyme pH Optima Overview EnzyBase12k is a specialized dataset containing 11,615 monomeric enzyme structures mapped to their experimentally determined pH optima. This dataset was curated to facilitate the development of machine learning models, specifically Equivariant Graph Neural Networks (EGNNs), capable of predicting biochemical properties directly from 3D protein geometry. The dataset was generated as part of the research: “Predicting Enzyme pH Optima from Structure Using Equivariant Graph Neural Networks” . --- Dataset Structure The dataset consists of two main parts: Structural Files: 11,615 files in `.cif` format (Crystallographic Information Framework). Metadata: `EnzyBase12k_metadata.csv` containing biological, structural, and experimental details. Dataset Files: 11,615 files in `.pqr` and also in `.pdb` Metadata Specification (`EnzyBase12k_metadata.csv`) ### Dataset Column Descriptions | Column Name | Description | Possible Values / Notes || :--- | :--- | :--- || **`uniprot_id`** | Unique identifier for each enzyme from the UniProt database. | e.g., `P00789` || **`ph_optimum`** | The optimal pH level for enzymatic activity (retrieved from UniProt/BRENDA). | Numeric value. || **`structure_source`** | Indicates if the structure is experimental or predicted. | `PDB` (Experimental)<br>`AF` (De novo predicted) || **`structure_mode`** | Specifies the method used to obtain the structure. | `AF2_downloaded` (AlphaFold DB)<br>`AF3` (AlphaFold 3 prediction)<br>`AF3_template` (AF3 using templates) || **`pdb_id_final`** | The identifier of the best-matching experimental PDB structure (if available). | e.g., `1ABC` || **`pdb_chain`** | The specific protein chain identifier corresponding to `pdb_id_final`. | e.g., `A`, `B` || **`ec_id`** | The Enzyme Commission number classifying the enzymatic reaction. | e.g., `1.1.1.1` || **`seq_cut_to_domain`** | Binary flag indicating if the sequence was truncated to a functional domain. | `1` (Yes, cut)<br>`0` (No, full length) || **`systematic_name`** | The systematic name of the enzyme. | e.g., *Alcohol dehydrogenase* || **`organism`** | The scientific name of the source organism. | e.g., *Homo sapiens* || **`seq_length`** | The total number of amino acids in the sequence used. | Integer value. || **`uniprot_seq_cut`** | The actual amino acid sequence used (full or truncated). | String sequence. | --- Licensing This dataset integrates data from multiple sources. Please adhere to the respective licenses: UniProt Data: [CC BY 4.0](https://www.uniprot.org/help/license)PDB Structures: [CC0 1.0 Universal](https://rcsb.org)AlphaFold2 Models: [CC BY 4.0](https://alphafold.ebi.ac.uk/)AlphaFold3 Models: [CC BY-NC-SA 4.0](https://github.com/google-deepmind/alphafold3) (Non-Commercial, ShareAlike) Metadata & Curation: The complete dataset curation and filtering process are done by Christian Clauß for the research project "Predicting Enzyme pH Optima from Structure Using Equivariant Graph Neural Networks". ---}

提供机构:
Zenodo
创建时间:
2026-01-28
二维码
社区交流群
二维码
科研交流群
商业服务