Companion dataset for FANTASIA v4.1.1: annotation examples and mouse benchmark results
收藏资源简介:
his record provides companion data for the FANTASIA v4.1.1 manuscript. The dataset contains benchmark and validation outputs generated with FANTASIA v4.1.1 for controlled protein-function annotation experiments based on protein language model embeddings. The companion archive includes two main case studies. First, it contains the PRUB1 annotation experiment, used as a representative annotation workflow. This includes the input protein FASTA file, run configuration, generated query embeddings, raw annotation outputs, summarized GO-term predictions, and TopGO-compatible export files. Second, it contains the MUSM_10090 mouse benchmark experiments, used to evaluate runtime, lookup behaviour, neighbourhood size, and filtering strategies. These benchmark runs include ProtT5 baseline analyses, lookup-only reuse of precomputed embeddings, distance-threshold and taxonomy-control runs, taxonomy plus MMseqs2 redundancy-masking runs, and k=5 model-comparison runs for ESM2, ESM3c, Ankh3-Large, ProstT5, and ProtT5. The data are intended to support reproducibility of the analyses reported in the manuscript and to document the structure of FANTASIA v4.1.1 outputs. The archive includes configuration files, logs or run-level metadata where available, raw per-query annotation tables, summary tables, identity-filter reports, TopGO export files, and helper scripts used for benchmark organization and comparison. This record is not a complete distribution of the FANTASIA software or reference databases. The FANTASIA v4.1.1 source code is available from the project GitHub repository, and the reference embedding databases are distributed separately through Zenodo. This companion dataset instead provides the manuscript-associated experimental outputs and benchmark material.



