遇见数据集

SwissProt Enzyme CDHIT60 Dataset

收藏
Zenodo2026-07-08 更新2026-08-02 收录
官方服务:

资源简介:

SwissProt Enzyme CDHIT60 Dataset 1 Dataset Overview To evaluate ConSite for enzyme active site annotation, we constructed a dataset from the Swiss-Prot section of UniProtKB by retaining enzymes with complete four-level EC annotations, at least one functional site annotation covering binding, catalytic, or other categories, and sequences of no more than 1,000 amino acids. Reaction SMILES were assigned through a hierarchical matching strategy that prioritizes curated UniProt–reaction mappings from RHEA, with EC-based matching as a fallback for enzymes lacking direct RHEA annotations. More than half of the dataset was matched through strict RHEA criteria requiring both UniProt entry presence and EC number concordance, ensuring more biologically grounded reaction assignments. The final dataset contains 102,080 active site annotation samples from approximately 95,000 proteins, covering 3,835 EC numbers across all seven enzyme classes, 73 subclasses, and 240 sub-subclasses. In total, over 700,000 functional residues are annotated, including 83.7% binding sites, 12.7% catalytic sites, and 3.6% other sites. Structural information was obtained from experimentally determined PDB structures or AlphaFold-predicted structures. 2 Data Splits & How to Use For evaluation, the dataset was split into training, validation, and test samples using CD-HIT clustering at a 60% sequence identity threshold to ensure that no test protein shares significant sequence similarity with training proteins. To support the hierarchical prototype alignment mechanism, we further applied an EC-aware stratified split ensuring that each test EC class has at least five training representatives, while we restricted validation and test sets to single-EC proteins to avoid ambiguity from multifunctional enzymes. `train.csv`: Training set containing 95,406 samples. `valid.csv`: Validation set containing 3,262 samples. `test.csv`: Test set containing 3,412 samples. 3 Data Columns Each CSV file contains the following columns: `rxn_smiles_single`: Reaction SMILES string representing the enzymatic reaction (substrate >> product). `ec`: Enzyme Commission (EC) number (e.g., 3.6.4.12). `pdb-id`: Associated PDB experimental structure ID(s), if applicable. `Uniprot ID`: UniProt accession ID. `Sequence`: The amino acid sequence of the enzyme. `site_labels`: A list of 0-indexed residue indices annotated as active sites on the sequence. `site_types`: The functional types corresponding to the `site_labels`, designated by integers (1 = binding site, 2 = catalytic site, 3 = other site).

提供机构:
Zenodo
创建时间:
2026-07-08
二维码
社区交流群
二维码
科研交流群
商业服务