FANTASIA V4 – LookUp Table – UniProt November 2025 – Experimental Evidence Code (Last Layer)
收藏资源简介:
FANTASIA V4 – LookUp Table UniProt November 2025 – Experimental Evidence Code (Layer 0 Only) Release: November 2025System: Protein Information System (PIS v3.1.0)Compatibility: FANTASIA V4.04 📖 Overview This PostgreSQL database backup uses the pgvector extension to store high-dimensional protein embeddings.It contains precomputed embeddings and functional annotations from the UniProt November 2025 release, restricted to entries with experimental evidence codes only. The lookup table was generated with PIS v3.1.0 (Protein Information System), an integrated platform for automated extraction, processing, and management of protein-related data.PIS consolidates information from UniProt, PDB, and GOA, enabling efficient retrieval of sequences, structures, and annotations. This release is designed for direct use with FANTASIA V4.04, an advanced pipeline for high-confidence functional annotation using Protein Language Models (PLMs).Unlike earlier releases, this dataset includes only layer 0 embeddings for each PLM model, providing a compact and efficient reference table for fast similarity search and GO term transfer. 🚫 Compatibility Notice This database is not compatible with versions of FANTASIA earlier than v4.0.4 andnot compatible with PIS versions earlier than v3.1.0. A tokenization inconsistency affecting the ProtT5-XL-UniRef50 model was corrected in this release.Because of this fix: ProtT5 embeddings produced with versions < FANTASIA v4.0.4 will not match those stored in this lookup table. Incompatibility only affects workflows that use the ProtT5 model. However, we highly recommend updating all components (FANTASIA, PIS, database) to ensure consistent behavior across all PLMs. This lookup table serves as a ready-to-use reference for large-scale protein function transfer: Loads layer-0 embeddings into memory Performs high-speed nearest-neighbor search in embedding space Transfers experimentally supported GO terms from annotated UniProt proteins It provides a stable, optimized, and fully curated base for reproducible annotation workflows within the FANTASIA ecosystem. Core Statistics UniProt accessions: 127,546 Protein records: 127,546 Unique sequences: 124,397 Includes 3,149 proteins sharing identical sequences (isoforms/redundancy) Experimental GO annotations: 627,932 Sequence redundancy: 2.47% Sequence Length Distribution (All Sequences) The 124,397 unique sequences exhibit a broad length distribution: Minimum: 3 aa Maximum: 35,375 aa Mean: 587.44 aa Q1: 262 aa Median: 431 aa Q3: 694 aa Missing ProtT5 Embeddings During embedding generation, the ProtT5-XL-UniRef50 model failed to process 150 sequences due to GPU memory limitations on an NVIDIA GeForce RTX 3090 Ti (24 GB VRAM). Although the dataset contains 127,546 protein records, only 150 proteins (≈0.12%) could not be embedded — a very small fraction of the total. These missing cases correspond exclusively to exceptionally long proteins, far beyond the typical length distribution: Minimum: 5,367 aa Maximum: 35,375 aa Mean: 9,715.53 aa Q1: 6,354.5 aa Median: 7,505.5 aa Q3: 8,930.5 aa All affected entries are provided in missing_prott5_embeddings_with_sequences.csv, and can be regenerated on systems with larger memory capacity. 🔬 Included GO Evidence Codes (Experimental Only) EXP — Inferred from Experiment IDA — Inferred from Direct Assay IPI — Inferred from Physical Interaction IMP — Inferred from Mutant Phenotype IGI — Inferred from Genetic Interaction IEP — Inferred from Expression Pattern TAS — Traceable Author Statement IC — Inferred by Curator
FANTASIA V4版——查找表 2025年11月版UniProt——仅实验证据编码(第0层) 发布版本:2025年11月 系统:蛋白质信息系统(Protein Information System,PIS v3.1.0) 兼容性要求:FANTASIA V4.04 📖 数据集概览 本PostgreSQL数据库备份采用pgvector扩展存储高维蛋白质嵌入向量,包含2025年11月版UniProt的预计算嵌入向量与功能注释,且仅纳入带有实验证据编码的条目。 本查找表由PIS v3.1.0(蛋白质信息系统,Protein Information System)生成,该平台是一款用于自动化提取、处理与管理蛋白质相关数据的集成化工具,整合了UniProt、蛋白质数据银行(Protein Data Bank,PDB)及基因本体注释(Gene Ontology Annotation,GOA)的信息,可实现序列、结构与注释的高效检索。 本版本专为配合FANTASIA V4.04使用而设计,后者是一款基于蛋白质语言模型(Protein Language Models,PLMs)实现高置信度功能注释的先进流程。与过往版本不同,本数据集仅包含各PLM模型的第0层嵌入向量,可为快速相似性搜索与基因本体(Gene Ontology,GO)术语迁移提供紧凑高效的参考表。 🚫 兼容性声明 本数据库无法与v4.0.4之前的FANTASIA版本及v3.1.0之前的PIS版本兼容。 本版本修复了影响ProtT5-XL-UniRef50模型的分词不一致问题,具体修复说明如下: - 采用FANTASIA v4.0.4之前版本生成的ProtT5嵌入向量,与本查找表中存储的嵌入向量无法匹配 - 该兼容性问题仅影响使用ProtT5模型的分析流程 尽管如此,我们强烈建议更新所有组件(FANTASIA、PIS及本数据库),以确保所有PLM模型的运行行为保持一致。 本查找表可作为大规模蛋白质功能迁移的即用型参考,具体功能包括: - 将第0层嵌入向量加载至内存 - 在嵌入空间中执行高速最近邻搜索 - 从已注释的UniProt蛋白质中迁移带有实验证据支持的GO术语 本数据集为FANTASIA生态系统内可复现的注释流程提供了稳定、优化且经过全面整理的基础支撑。 核心统计数据 - UniProt登录号数量:127546个 - 蛋白质记录数:127546条 - 唯一序列数:124397条 - 包含3149条序列完全一致的蛋白质(即同工型/冗余序列) - 实验支持的GO注释数:627932条 - 序列冗余率:2.47% 序列长度分布(所有序列) 124397条唯一序列的长度分布跨度广泛: - 最小值:3个氨基酸残基(aa) - 最大值:35375个氨基酸残基 - 平均值:587.44个氨基酸残基 - 四分位数Q1:262个氨基酸残基 - 中位数:431个氨基酸残基 - 四分位数Q3:694个氨基酸残基 缺失的ProtT5嵌入向量 在嵌入向量生成过程中,由于NVIDIA GeForce RTX 3090 Ti(24GB显存)的GPU内存限制,ProtT5-XL-UniRef50模型无法处理150条序列。 尽管本数据集共包含127546条蛋白质记录,但仅有150条蛋白质(约占总数的0.12%)未能生成嵌入向量,占比极低。 这些缺失的序列均为超长蛋白质,其长度远超常规分布范围: - 最小值:5367个氨基酸残基 - 最大值:35375个氨基酸残基 - 平均值:9715.53个氨基酸残基 - 四分位数Q1:6354.5个氨基酸残基 - 中位数:7505.5个氨基酸残基 - 四分位数Q3:8930.5个氨基酸残基 所有受影响的条目均已收录于missing_prott5_embeddings_with_sequences.csv文件中,可在显存容量更大的系统中重新生成嵌入向量。 🔬 纳入的实验类GO证据编码 EXP — 实验推断(Inferred from Experiment) IDA — 直接实验测定推断(Inferred from Direct Assay) IPI — 物理相互作用推断(Inferred from Physical Interaction) IMP — 突变体表型推断(Inferred from Mutant Phenotype) IGI — 遗传相互作用推断(Inferred from Genetic Interaction) IEP — 表达模式推断(Inferred from Expression Pattern) TAS — 可追溯作者声明(Traceable Author Statement) IC — 馆长推断(Inferred by Curator)



