FANTASIA V4 – LookUp Table – UniProt November 2025 – Experimental Evidence Code (Last Layer)
收藏资源简介:
FANTASIA V4 – LookUp Table UniProt November 2025 – Experimental Evidence Code (Layer 0 Only) Release: November 2025System: Protein Information System (PIS v3.1.0)Compatibility: FANTASIA V4.04 📖 Overview This PostgreSQL database backup uses the pgvector extension to store high-dimensional protein embeddings.It contains precomputed embeddings and functional annotations from the UniProt November 2025 release, restricted to entries with experimental evidence codes only. The lookup table was generated with PIS v3.1.0 (Protein Information System), an integrated platform for automated extraction, processing, and management of protein-related data.PIS consolidates information from UniProt, PDB, and GOA, enabling efficient retrieval of sequences, structures, and annotations. This release is designed for direct use with FANTASIA V4.04, an advanced pipeline for high-confidence functional annotation using Protein Language Models (PLMs).Unlike earlier releases, this dataset includes only layer 0 embeddings for each PLM model, providing a compact and efficient reference table for fast similarity search and GO term transfer. 🚫 Compatibility Notice This database is not compatible with versions of FANTASIA earlier than v4.0.4 andnot compatible with PIS versions earlier than v3.1.0. A tokenization inconsistency affecting the ProtT5-XL-UniRef50 model was corrected in this release.Because of this fix: ProtT5 embeddings produced with versions < FANTASIA v4.0.4 will not match those stored in this lookup table. Incompatibility only affects workflows that use the ProtT5 model. However, we highly recommend updating all components (FANTASIA, PIS, database) to ensure consistent behavior across all PLMs. This lookup table serves as a ready-to-use reference for large-scale protein function transfer: Loads layer-0 embeddings into memory Performs high-speed nearest-neighbor search in embedding space Transfers experimentally supported GO terms from annotated UniProt proteins It provides a stable, optimized, and fully curated base for reproducible annotation workflows within the FANTASIA ecosystem. Core Statistics UniProt accessions: 127,546 Protein records: 127,546 Unique sequences: 124,397 Includes 3,149 proteins sharing identical sequences (isoforms/redundancy) Experimental GO annotations: 627,932 Sequence redundancy: 2.47% Sequence Length Distribution (All Sequences) The 124,397 unique sequences exhibit a broad length distribution: Minimum: 3 aa Maximum: 35,375 aa Mean: 587.44 aa Q1: 262 aa Median: 431 aa Q3: 694 aa Missing ProtT5 Embeddings During embedding generation, the ProtT5-XL-UniRef50 model failed to process 150 sequences due to GPU memory limitations on an NVIDIA GeForce RTX 3090 Ti (24 GB VRAM). Although the dataset contains 127,546 protein records, only 150 proteins (≈0.12%) could not be embedded — a very small fraction of the total. These missing cases correspond exclusively to exceptionally long proteins, far beyond the typical length distribution: Minimum: 5,367 aa Maximum: 35,375 aa Mean: 9,715.53 aa Q1: 6,354.5 aa Median: 7,505.5 aa Q3: 8,930.5 aa All affected entries are provided in missing_prott5_embeddings_with_sequences.csv, and can be regenerated on systems with larger memory capacity. 🔬 Included GO Evidence Codes (Experimental Only) EXP — Inferred from Experiment IDA — Inferred from Direct Assay IPI — Inferred from Physical Interaction IMP — Inferred from Mutant Phenotype IGI — Inferred from Genetic Interaction IEP — Inferred from Expression Pattern TAS — Traceable Author Statement IC — Inferred by Curator



