This collection contains the benchmark data used for benchmarking text extraction tools. The data conatins : - List of documents - Ground truth data for each document - Additional
aThe first isolate listed for each species is the approved or suggested prototype strain for that species. bValues in parentheses for each isolate indicate which of its genome segments (by size rank s