Evolutionary fingerprinting of virus proteins
收藏资源简介:
This is a repository of data associated with a project ("Evolutionary fingerprinting of virus proteins") with the objective of comparing the pattern of natural selection (via comparative dN/dS analysis) acting on different proteins encoded by the genomes of different virus species, and finding common determinants of those patterns. All data are derived from NCBI Genbank, with original sources and contributors identified by accession numbers. step1.tar.gz - a gzip-compressed TAR of all step1 plain text (.txt) files containing Genbank accession lists step3.tar.gz - a gzip-compressed TAR of all step 3 FASTA files (codon-aligned and manually revised) step3_md5.txt - MD5 checksums of individual step 3 FASTA files step4.tar.gz - a gzip-compressed TAR of all step 4 FASTA files (with removal of sites affected by gene overlap independently of step 7) step4_md5.txt - MD5 checksums of individual step 4 FASTA files step4_json.tar.gz - a gzip-compressed TAR of all step4 JSON files step5.tar.gz - a gzip-compressed TAR of all step 5 FASTA files (down-sampling to normalize tree lengths) step5_md5.txt - MD5 checksums of individual step 5 FASTA files step7.tar.gz - a gzip-compressed TAR of all step 7 FASTA files (removal of sites affected by gene overlap) step7_md5.txt - MD5 checksums of individual step 7 FASTA files L50.RData - a binary file containing two R objects: (1) a matrix of Wasserstein distances (`wmat`) and (2) a data frame containing metadata; these objects correspond to replicate samples of 50 codons from each alignment except for those below this target length (these are represented by one entry only). L100.RData - a binary file containing two R objects: (1) a matrix of Wasserstein distances (`wmat`) and (2) a data frame containing metadata; these objects correspond to replicate samples of 100 codons from each alignment, excluding alignments shorter than 100 codons.



