遇见数据集

HiFEN dataset

收藏
Zenodo2026-09-28 更新2026-10-01 收录
官方服务:

资源简介:

This dataset was prepared for the development and evaluation of HiFEN (Hierarchical Functional Enzyme Network), a deep-learning framework for hierarchical enzyme function prediction. HiFEN predicts one or more Enzyme Commission (EC) labels at each of the four hierarchical levels, from EC1 to EC4. The dataset contains 166,018 protein records with amino acid sequences, hierarchical EC annotations, clustering information, and residue-level functional annotations. It supports both single-label and multi-label enzyme classification and is suitable for training, validation, and benchmarking of hierarchical protein function prediction models. The dataset is used by the HiFEN project. The corresponding source code and usage instructions are available at: https://github.com/LycrsLOL/HiFEN Dataset contents The dataset is distributed as a CSV file containing the following columns: uniprot_id: UniProt accession or protein identifier. pdb_id: associated Protein Data Bank identifier, when available. ec_numbers: complete set of EC annotations associated with the protein. sequence: amino acid sequence of the protein. homology_cluster: homology-based cluster assignment used to control sequence similarity across data partitions. sequence_cluster: sequence-based cluster assignment. active_sites: annotated active-site residues. catalytic_sites: annotated catalytic residues. binding_sites: annotated ligand- or substrate-binding residues. domain_ranges: annotated protein-domain residue ranges. motif_ranges: annotated functional-motif residue ranges. ec_numbers_supervised: EC labels used as supervised learning targets. ec_numbers_ignored: EC annotations excluded from supervised targets, for example because they do not satisfy the applicable annotation or data-selection criteria. Multiple values within a field are serialized in the CSV representation. Empty fields indicate that the corresponding annotation or identifier is unavailable. Task definition EC prediction is formulated hierarchically at four levels: EC1: top-level enzyme class EC2: subclass EC3: sub-subclass EC4: complete EC annotation A protein may have one or multiple EC annotations. Consequently, the dataset can be used for both single-label and multi-label evaluation. For multi-label experiments, predictions and metrics should be calculated independently at EC1, EC2, EC3, and EC4. The provided homology and sequence cluster fields can be used to construct data partitions that reduce sequence-level information leakage between training, validation, and test sets. Users should keep all members of the same relevant cluster within a single partition when creating homology-aware splits. Recommended evaluation For multi-label evaluation, recommended metrics include: Micro-precision Micro-recall Micro-F1 Macro-F1 Observed-class macro-F1 Micro-AUPRC Macro-AUPRC Micro-AUROC Macro-AUROC F-max Matthews correlation coefficient Hamming loss Subset accuracy Label-ranking average precision Class-specific decision thresholds should be calibrated exclusively on a validation set and fixed before evaluation on the test set. The test set should not be used for threshold calibration, model selection, or hyperparameter optimization. For single-label comparisons, only records and labels compatible with the corresponding single-label task definition should be used. Results from single-label methods should not be directly compared with multi-label metrics. Scope and exclusions This release provides protein sequences, EC annotations, clustering information, and residue-level functional annotations required for model development. It does not include: Predicted or experimentally determined structure files Precomputed protein embeddings Graph caches Model checkpoints Training logs Evaluation outputs Private filesystem paths or computational-environment information Any structural features, sequence embeddings, or graph representations required by a downstream model must be generated separately by the user. Data usage This dataset is intended for academic research involving: Hierarchical enzyme function prediction Multi-label protein classification Enzyme annotation Protein representation learning Homology-aware model evaluation Integration of sequence, structure, and residue-level functional information Benchmarking of enzyme function prediction methods Users should document their splitting strategy, EC vocabulary, label-processing procedure, decision-threshold calibration method, and evaluation metrics to support reproducibility and fair comparison. Related software This dataset is used by: HiFEN — Hierarchical Functional Enzyme Networkhttps://github.com/LycrsLOL/HiFEN For Zenodo’s Related identifiers field, the HiFEN repository can be registered as: Identifier: https://github.com/LycrsLOL/HiFEN Relation: Is used by Resource type: Software

提供机构:
Zenodo
创建时间:
2026-09-28
二维码
社区交流群
二维码
科研交流群
商业服务