Identification of protein domains is a key step for understanding protein function. Hidden Markov Models (HMMs) have proved to be a powerful tool for this task. The Pfam database notably provides a la
The contents of this dataset were created de novo starting from RNA-seq data downloaded from Gene Expression Omnibus (GEO), with a pipeline specifically written to generate results starting with seque
Additional file 2. Source code for constructing benchmarks: Source code in an archive format, using tar and bzip2, for users to generate their own benchmarks for any genomic text in FASTA format. This