遇见数据集

A comprehensive dataset of homologous and non-homologous isofunctional enzymes across the tree of life

收藏
Zenodo2026-04-07 更新2026-05-26 收录
官方服务:

资源简介:

A comprehensive dataset of homologous and non-homologous isofunctional enzymes across the tree of life Background Convergent evolution, the independent emergence of similar traits, is increasingly recognized as a pervasive force shaping molecular and metabolic diversity. A striking manifestation of convergence at the molecular level is represented by non-homologous isofunctional enzymes (NISE), distinct proteins with no detectable common ancestry that catalyze identical biochemical reactions. Despite their conceptual and practical relevance, NISE are often treated as exceptional cases, and no large-scale, systematically curated resource has been available to explore their distribution and properties across all domains of life. Data description Here we present a curated dataset of homologous and non-homologous isofunctional enzymes (HISE and NISE) derived from UniProtKB release 2025_01, encompassing both reviewed (Swiss-Prot) and unreviewed (TrEMBL) entries. Using Enzyme Commission (EC) numbers to define catalytic equivalence and SUPERFAMILY (SCOP structure superfamily) annotations to infer evolutionary relationships, we implemented a transparent and reproducible pipeline to classify enzymes into homologous and non-homologous functional groups. The dataset comprises over 200,000 Swiss-Prot and 27 million TrEMBL enzymes with complete EC and SUPERFAMILY annotations, organized by domain of life, enzyme class, and structural domain composition. Multiple output files, including presence/absence matrices, clustered enzyme groups, phyloprofiles, and full annotation tables, are provided to facilitate downstream evolutionary, functional, and comparative analyses. This resource offers a global view of molecular convergence and divergence in enzymatic functions, highlighting the widespread nature of NISE across taxa and enzyme classes. It provides a foundation for studying metabolic evolution, functional redundancy, drug target discovery, and the evolutionary constraints shaping biochemical solutions. Methodology Identification and classification of homologous and non-homologous isofunctional enzymes Data source and retrieval Protein sequence and annotation data were obtained from UniProtKB (release 2025_01). All UniProtKB entries annotated with catalytic activity were retrieved, comprising 257,171 reviewed (Swiss-Prot) and 33,961,726 unreviewed (TrEMBL) sequences, spanning 8,562 Swiss-Prot and 1,050,770 TrEMBL organisms, including viruses. Entries were downloaded in tab-separated (TSV) format and included the following fields: UniProt accession, entry name, protein names, gene names, organism name, Enzyme Commission (EC) number, SUPERFAMILY cross-references, and complete taxonomic lineage. Only proteins annotated with a single, complete EC number and at least one SUPERFAMILY (SUPFAM) domain assignment were retained for further analysis. Proteins annotated with multiple EC numbers (promiscuous enzymes), incomplete EC numbers, or lacking SUPERFAMILY annotations were excluded from downstream analyses and recorded separately in a log file (NISE.log). After filtering, the curated dataset comprised 203,921 Swiss-Prot and 27,371,520 TrEMBL enzyme sequences. Construction of SUPERFAMILY presence/absence matrices For each distinct EC number identified in the filtered dataset (4,575 in Swiss-Prot and 2,135 in TrEMBL), a presence/absence matrix of SUPERFAMILY annotations was constructed. In these matrices, rows correspond to enzyme sequences and columns correspond to SUPERFAMILY domains, with binary values indicating the presence or absence of a given domain. These matrices were exported to the file NISE.mtx and constituted the basis for subsequent classification of homologous and non-homologous isofunctional enzymes. Identification of homologous isofunctional enzymes (HISE) EC numbers were classified as encoding homologous isofunctional enzymes (HISE) when all associated enzyme sequences shared identical SUPERFAMILY profiles: 1. Single-domain HISE (HISE.bnf): EC numbers for which all enzymes displayed exactly one SUPFAM annotation and shared the same SUPFAM were collected. For each EC, the number of clusters and enzyme sequences was recorded and written to HISE.bnf. This category comprised 2,090 Swiss-Prot and 1,271 TrEMBL EC numbers, encompassing 101,303 Swiss-Prot and 8,733,696 TrEMBL sequences. 2. Multi-domain HISE (HISE.multi): EC numbers for which all enzymes shared identical multiple SUPFAM annotations (i.e. identical multi-domain architectures) were collected and written to HISE.multi. This category included 510 Swiss-Prot and 71 TrEMBL EC numbers, comprising 14,631 Swiss-Prot and 58,588 TrEMBL sequences. 3. Singleton HISE (HISE.sgl): EC numbers represented by a single enzyme sequence, displaying either one or multiple SUPFAM annotations, were classified as singleton HISE and written to HISE.sgl. This category comprised 1,498 Swiss-Prot and 92 TrEMBL EC numbers. Identification of non-homologous isofunctional enzymes (NISE) Non-homologous isofunctional enzymes (NISE) were identified as enzymes sharing the same EC number but associated with two or more distinct SUPERFAMILY profiles, indicating independent evolutionary origins. 4. SUPFAM-based clustering (NISE.clust): For each EC number, enzymes were grouped into clusters based on identical SUPFAM profiles. EC numbers comprising at least two distinct SUPFAM profiles were retained, and cluster compositions were written to NISE.clust. This resulted in 914 Swiss-Prot and 1,689 TrEMBL EC numbers, encompassing 112,093 Swiss-Prot and 26,977,106 TrEMBL sequences distributed across 3,141 and 54,111 clusters, respectively. 5. Multi-domain NISE (NISE.multi): Enzymes displaying more than one SUPFAM annotation per sequence were classified as multi-domain NISE. EC numbers containing such enzymes were collected and written to NISE.multi, comprising 355 Swiss-Prot and 1,456 TrEMBL EC numbers and a total of 37,613 Swiss-Prot and 10,303,985 TrEMBL sequences. 6. Monomeric enzymes retained for NISE analysis: Enzymes displaying a single SUPFAM annotation per sequence were retained as monomeric enzymes for downstream NISE analyses. This set included 559 Swiss-Prot and 233 TrEMBL EC numbers, comprising 74,480 Swiss-Prot and 16,673,121 TrEMBL sequences. 7. Single-domain NISE (NISE.bnf): EC numbers containing at least two enzymes sharing distinct SUPFAM annotations (each enzyme with a single SUPFAM) were classified as single-domain non-homologous isofunctional enzymes and written to NISE.bnf. This category comprised 338 Swiss-Prot and 540 TrEMBL EC numbers, encompassing 32,895 Swiss-Prot and 8,245,833 TrEMBL sequences grouped into 846 and 4,202 clusters, respectively. Logging of excluded entries All enzymes excluded from the main analyses were recorded in NISE.log, together with full UniProtKB annotation. This file includes: * Enzymes without EC annotation (5,927 Swiss-Prot; 2,105,913 TrEMBL),* Enzymes without SUPERFAMILY annotation (17,696 Swiss-Prot; 2,231,328 TrEMBL),* Enzymes without complete EC numbers (16,046 Swiss-Prot; 1,410,797 TrEMBL),* Promiscuous enzymes annotated with multiple EC numbers (17,771 Swiss-Prot; 1,267,621 TrEMBL). This protocol provides a transparent, reproducible framework for the large-scale identification and classification of homologous and non-homologous isofunctional enzymes and can be readily adapted to future UniProtKB releases or alternative domain classification systems. Description of the data files File name Description Format Content SwissProt_NISE.tsv.gz Full annotation table for non-homologous isofunctional enzymes TSV UniProtKB accession, protein and gene names, EC number, SUPERFAMILY annotation(s), organism name, and taxonomic ranks from species to domain SwissProt_NISE.mtx.gz SUPERFAMILY presence/absence matrix by EC number TSV Binary matrix indicating SUPERFAMILY domains associated with enzymes sharing the same EC SwissProt_NISE.clust.gz Clustered non-homologous isofunctional enzymes TXT EC numbers, SUPERFAMILY profiles (mono and multi-domain architectures), and enzyme identifiers SwissProt_NISE.bnf.gz Non-homologous isofunctional monomeric enzymes dataset TXT EC numbers, SUPERFAMILY identifiers, and enzyme identifiers for SUPERFAMILY mono-domain architectures enzymes SwissProt_NISE.multi.gz Non-homologous isofunctional multimeric enzymes dataset TXT EC numbers, SUPERFAMILY identifiers, and enzyme identifiers for SUPERFAMILY multi-domain architectures enzymes SwissProt_NISE.log.gz Excluded entries (missing EC/SUPFAM, incomplete or promiscuous EC) TSV UniProtKB accession, protein and gene names, EC number, SUPERFAMILY annotation(s), organism name, and taxonomic ranks from domain to species SwissProt_HISE.tsv.gz Full annotation table for homologous isofunctional enzymes TSV Same fields as NISE.tsv for enzymes classified as homologous SwissProt_HISE.bnf.gz Homologous isofunctional enzymes (monomeric) TXT Same fields as NISE.bnf for enzymes classified as homologous SwissProt_HISE.multi.gz Homologous isofunctional enzymes (multimeric) TXT Same fields as NISE.multi for enzymes classified as homologous SwissProt_HISE.sgl.gz Singleton homologous enzymes TXT EC numbers, SUPERFAMILY identifiers, and enzyme identifiers for EC numbers represented by a single enzyme SwissProt_PhyloProfiles.tar.gz Input matrices files (*.phyprof) for PhyloProfile software TSV Formatted presence/absence matrices linking SUPERFAMILY domains, taxa, and enzyme identifiers TrEMBL_NISE.tsv.gz Full annotation table for non-homologous isofunctional enzymes TSV UniProtKB accession, protein and gene names, EC number, SUPERFAMILY annotation(s), organism name, and taxonomic ranks from species to domain TrEMBL_NISE.mtx.gz SUPERFAMILY presence/absence matrix by EC number TSV Binary matrix indicating SUPERFAMILY domains associated with enzymes sharing the same EC TrEMBL_NISE.clust.gz Clustered non-homologous isofunctional enzymes TXT EC numbers, SUPERFAMILY profiles (mono and multi-domain architectures), and enzyme identifiers TrEMBL_NISE.bnf.gz Non-homologous isofunctional monomeric enzymes dataset TXT EC numbers, SUPERFAMILY identifiers, and enzyme identifiers for SUPERFAMILY mono-domain architectures enzymes TrEMBL_NISE.multi.gz Non-homologous isofunctional multimeric enzymes dataset TXT EC numbers, SUPERFAMILY identifiers, and enzyme identifiers for SUPERFAMILY multi-domain architectures enzymes TrEMBL_NISE.log.gz Excluded entries (missing EC/SUPFAM, incomplete or promiscuous EC) TSV UniProtKB accession, protein and gene names, EC number, SUPERFAMILY annotation(s), organism name, and taxonomic ranks from domain to species TrEMBL_HISE.tsv.gz Full annotation table for homologous isofunctional enzymes TSV Same fields as NISE.tsv for enzymes classified as homologous TrEMBL_HISE.bnf.gz Homologous isofunctional enzymes (monomeric) TXT Same fields as NISE.bnf for enzymes classified as homologous TrEMBL_HISE.multi.gz Homologous isofunctional enzymes (multimeric) TXT Same fields as NISE.multi for enzymes classified as homologous TrEMBL_HISE.sgl.gz Singleton homologous enzymes TXT EC numbers, SUPERFAMILY identifiers, and enzyme identifiers for EC numbers represented by a single enzyme TrEMBL_PhyloProfiles.tar.gz Input matrices files (*.phyprof) for PhyloProfile software TSV Formatted presence/absence matrices linking SUPERFAMILY domains, taxa, and enzyme identifiers Stats.txt NISE and HISE global statistics TXT NISE/HISE counts by domain of life and enzyme class SwissProt.png Proportional distribution of NISE and HISE in SWISS-Prot PNG Histograms of NISE and HISE counts across domains of life and enzyme classes in Swiss-Prot data TrEMBL.png Proportional distribution of NISE and HISE in TrEMBL PNG Histograms of NISE and HISE counts across domains of life and enzyme classes in TrEMBL data Flowchart.png Workflow of the methodology PNG Schematic representation of the of the pipeline, illustrating data filtering, classification, and output generation Applications All data are structured, parseable flat files in TXT or TSV format, except figures. The dataset enables large-scale analyses of enzymatic convergence and divergence, supporting comparative and evolutionary genomics studies, providing a resource for identifying alternative biochemical solutions to the same metabolic reactions. Profiles representing the taxonomic distribution of NISE and HISE across domains, kingdoms, and lower taxonomic ranks for each EC numer (*.phyprof) - packed and compressed in files SwissProt/TrEMBL_PhyloProfiles.tar.gz - are compatible with the PhyloProfile software <https://github.com/BIONF/PhyloProfile>, a comparative evolutionary framework, in which: enzyme families distributed across domains of life (Archaea, Bacteria, Eukarya, Viruses) can be visualized; patterns of conservation, gain, loss, or convergence can be detected; domain architectures correlating with specific lineages, metabolic capabilities, or evolutionary events can be investigated. Data was derived from the following source * UniProtKB <https://www.uniprot.org/>

提供机构:
Zenodo
创建时间:
2025-12-23
二维码
社区交流群
二维码
科研交流群
商业服务