遇见数据集

10-Fold Cross-Validation Area-Under-the-Curve scores (mean ± st.dev.) of protein representation methods in the inference tasks. Hist-8000 outperforms SoT in seven out of tasks 1-8. Hist-8000 consists of the conceptually simpler BoW approach [13], in contrast to SoT which requires SSL pre-training (word2vec) on a large protein sequence dataset [17]. Best-performing methods in bold. See section ‘Protein inference problems’ for task data sources. Gram-negative, Gram-positive and Archaea datasets are each used for subcellular localisation prediction from protein sequence. Gram-neg.: Gram-negative bacteria, Gram-pos.: Gram-positive bacteria, #Proteins (pos.+neg.): number of proteins (positive+negative), Hist-8000: Histogram-8000, BoW: Bag-of-Word

收藏
NIAID Data Ecosystem2026-05-02 收录
官方服务:

资源简介:

10-Fold Cross-Validation Area-Under-the-Curve scores (mean ± st.dev.) of protein representation methods in the inference tasks. Hist-8000 outperforms SoT in seven out of tasks 1-8. Hist-8000 consists of the conceptually simpler BoW approach [13], in contrast to SoT which requires SSL pre-training (word2vec) on a large protein sequence dataset [17]. Best-performing methods in bold. See section ‘Protein inference problems’ for task data sources. Gram-negative, Gram-positive and Archaea datasets are each used for subcellular localisation prediction from protein sequence. Gram-neg.: Gram-negative bacteria, Gram-pos.: Gram-positive bacteria, #Proteins (pos.+neg.): number of proteins (positive+negative), Hist-8000: Histogram-8000, BoW: Bag-of-Word

创建时间:
2025-08-06
二维码
社区交流群
二维码
科研交流群
商业服务