Supporting data for "A close look at protein function prediction evaluation protocols".

Name: Supporting data for "A close look at protein function prediction evaluation protocols".
Creator: GigaScience Database
Published: 2025-05-26 17:25:22
License: 暂无描述

DataCite Commons2025-05-26 更新2025-04-15 收录

下载链接：

http://gigadb.org/dataset/100153

下载链接

链接失效反馈

官方服务：

资源简介：

The recently held Critical Assessment of Functional Annotation challenge (CAFA2) required its participants to submit predictions for a large number of target proteins regardless of whether they have previous annotations or not. This is in contrast to the original CAFA challenge in which participants were asked to submit predictions for proteins with no existing annotations. The CAFA2 task is more realistic, in that it more closely mimics the accumulation of annotations over time. In this study we compare these tasks in terms of their difficulty, and determine if cross-validation provides a good estimate of performance. The CAFA2 task is a combination of two sub-tasks: making predictions on annotated proteins and making predictions on previously unannotated proteins. In this study we analyze the performance of several function prediction methods in these two scenarios. Our results show that several methods (GOstruct, binary SVMs, and guilt by association) find it hard to achieve the same level of accuracy on these two tasks compared to cross-validation, and that predicting novel annotations for previously annotated proteins is a harder problem than predicting annotations for uncharacterized proteins. We also find that different methods have different performance characteristics in these tasks, and that cross-validation is not adequate at estimating performance and ranking methods.

提供机构：

GigaScience Database

创建时间：

2015-08-24

5,000+

优质数据集

54 个

任务类型

进入经典数据集