Published protein sequence sets for the evaluation of bioinformatics tools
收藏资源简介:
A collection of protein sequence sets drawn from published studies and public sequence databases, assembled for the evaluation of bioinformatics tools. Each set corresponds to a protein family or superfamily that has been analysed in the published literature, and is accompanied where available by the reference annotations and classifications published alongside it. Sequence identifiers are preserved in a form that allows the accompanying annotation tables to be joined directly onto tool output. The sets are archived here so that evaluations can be repeated on identical inputs. This is not a convenience: several are snapshots whose underlying database records have since been revised or withdrawn, and cannot be reconstructed from a current query. Provenance for each set — its source, the published work it derives from, and its known limitations — is documented in the accompanying manifest.



