10-fold Cross-Validation Area-Under-the-Curve scores (mean ± st.dev.) of ProtT5 vs Histogram-8000 protein representation methods in the inference tasks. Hist-8000 matches ProtT5 in four out of seven tasks compared (tasks 1, 3-8). Hist-8000 consists of the conceptually simpler BoW approach [13], in contrast to ProtT5 which requires SSL pre-training of a transformer T5 model with three billion parameters on a dataset of 45 million sequences [35]. Best-performing methods in bold. See section ‘Protein inference problems’ for task data sources. Hist-8000: Histogram-8000, BoW: Bag-of-Words, SoT: Sum-of-learnt-Trigrams, SSL: Self-Supervised Learning, VFs: Virulence Factors, Gram-pos: Gram-positive, Gram-neg: Gram-negative, st.dev: standard deviatio
收藏资源简介:
10-fold Cross-Validation Area-Under-the-Curve scores (mean ± st.dev.) of ProtT5 vs Histogram-8000 protein representation methods in the inference tasks. Hist-8000 matches ProtT5 in four out of seven tasks compared (tasks 1, 3-8). Hist-8000 consists of the conceptually simpler BoW approach [13], in contrast to ProtT5 which requires SSL pre-training of a transformer T5 model with three billion parameters on a dataset of 45 million sequences [35]. Best-performing methods in bold. See section ‘Protein inference problems’ for task data sources. Hist-8000: Histogram-8000, BoW: Bag-of-Words, SoT: Sum-of-learnt-Trigrams, SSL: Self-Supervised Learning, VFs: Virulence Factors, Gram-pos: Gram-positive, Gram-neg: Gram-negative, st.dev: standard deviatio



