10-Fold Cross-Validation Area-Under-the-Curve scores (mean ± st.dev.) of protein representation methods in the inference tasks. Hist-8000 outperforms SoT in seven out of tasks 1-8. Hist-8000 consists of the conceptually simpler BoW approach [13], in contrast to SoT which requires SSL pre-training (word2vec) on a large protein sequence dataset [17]. Best-performing methods in bold. See section ‘Protein inference problems’ for task data sources. Gram-negative, Gram-positive and Archaea datasets are each used for subcellular localisation prediction from protein sequence. Gram-neg.: Gram-negative bacteria, Gram-pos.: Gram-positive bacteria, #Proteins (pos.+neg.): number of proteins (positive+negative), Hist-8000: Histogram-8000, BoW: Bag-of-Word
收藏资源简介:
10-Fold Cross-Validation Area-Under-the-Curve scores (mean ± st.dev.) of protein representation methods in the inference tasks. Hist-8000 outperforms SoT in seven out of tasks 1-8. Hist-8000 consists of the conceptually simpler BoW approach [13], in contrast to SoT which requires SSL pre-training (word2vec) on a large protein sequence dataset [17]. Best-performing methods in bold. See section ‘Protein inference problems’ for task data sources. Gram-negative, Gram-positive and Archaea datasets are each used for subcellular localisation prediction from protein sequence. Gram-neg.: Gram-negative bacteria, Gram-pos.: Gram-positive bacteria, #Proteins (pos.+neg.): number of proteins (positive+negative), Hist-8000: Histogram-8000, BoW: Bag-of-Word



