Prospects for a sequence-based taxonomy of influenza A virus subtypes
收藏资源简介:
This dataset comprises the multiple sequence alignments (*.fasta) and maximum likelihood phylogenies (Newick tree strings, *.nwk) for all available protein sequences corresponding to the eight genome segments of influenza A virus from the NCBI Genbank database. Each sequence is labelled with the Genbank accession number (e.g., "CY103884"), WHO strain identifier ("A/little yellow-shouldered bat/Guatemala/164/2009"), subtype label ("H17N10"), host species ("Sturnira lilium; gender M"), sampling location ("Guatemala: El Jobo"), and sample collection date ("May-2009"). These fields are separated by underscore characters. These data are provided under a Creative Commons license in support of a manuscript in progress, "Prospects for a sequence-based taxonomy of influenza A virus subtypes".



