Cross-Species Comparative Analysis of Influenza A Virus Proteins (HA, NA, NP, PA, PB1, PB2, M1, M2, NS1, NS2)
收藏资源简介:
This dataset comprises a comprehensive collection of Influenza A virus protein sequences across multiple host species. It includes hemagglutinin (HA), neuraminidase (NA), nucleoprotein (NP), polymerase complex (PA, PB1, PB2), matrix proteins (M1, M2), and nonstructural proteins (NS1, NS2). The data enable comparative genomic and proteomic exploration of host-adaptive and evolutionary patterns. These have been organized by protein type (HA, NA, NP, PA, PB1, PB2, M1, M2, NS1, NS2). Each protein folder contains: Individual FASTA files for each sequence. QC reports (.txt) with sequence metrics. Metadata (.tsv) including sequence ID, full header, description, length, GC content, and ambiguous bases. Panel plots (.png) and count files (.txt) visualizing the distribution of sequences by city, year, and country. This dataset is for my journal article. Please cite the dataset if you use it. Global Summary of H5N1 Proteins Metadata: Top Countries/Cities: Unknown: 84530 New York: 16 Guangdong: 16 Hong Kong: 15 Shanghai: 15 California: 15 Puerto Rico: 15 Korea: 14 Ann Arbor: 13 Oklahoma: 13 Lee: 12 Top Years: Unknown: 42221 2004: 16 1996: 16 2013: 15 2009: 15 1934: 15 1968: 14 2011: 13 1940: 12 Data processed by: TahirHB@Hotmail.Com Note: Due to initial keyword-based parsing, a subset of entries includes non-avian Influenza A protein records. These have been retained for cross-species comparison and phylogenetic exploration. Note: Dataset expanded from H5N1-specific proteins to include additional Influenza A virus sequences detected across multiple host species and subtypes. Enables broader comparative and evolutionary analysis of HA, NA, NP, PA, PB1, PB2, M1, M2, NS1, and NS2 proteins.




