DRGV: metadata, genomes, proteins
收藏资源简介:
This record contains the metadata tables, genome sequences, protein catalogs, and functional annotation of DRGV. The Dog Reference Gut Microbiome (DRGM), the prokaryotic catalog from which the predicted viral hosts are drawn, is deposited as separate records. Viral metadata DRGV_genome_metadata.tsv : per-genome information for all 20,875 viral genomes - genome ID, vOTU assignment, source BioProject and sample, genome length, sequence type, CheckV quality, confidence method and result, completeness, and representative flag DRGV_vOTU_metadata.tsv : cluster-level information for the 12,125 vOTUs - representative genome, length, source BioProject and sample, CheckV quality, confidence method and result, completeness, cluster size, ICTV taxonomy from realm to genus, human-virome match rank, predicted DRGM host and host rank, and whether the vOTU is in the terminase-large-subunit tree Viral representative genomes DRGV_representative_genomes.tar.gz : genome sequences (FASTA) for the 12,125 DRGV vOTU representative genomes, one per vOTU All viral genomes DRGV_all_genomes.tar.gz : genome sequences (FASTA) for all 20,875 viral genomes used to define the DRGV vOTUs Viral proteins Protein catalogs at seven levels of redundancy, in the same layout as the DRGM protein catalogs. "unique" holds the sequences left after exact-duplicate removal; the others are clustered at the stated amino-acid identity. Clustering at 100% identity is therefore not the same as the unique set, since the sequences in a family need not be of equal length. DRGV_unique_proteins.tar.gz : unique protein sequences after redundancy removal DRGV_100_proteins.tar.gz : protein families at 100% identity DRGV_95_proteins.tar.gz : protein families at 95% identity DRGV_90_proteins.tar.gz : protein families at 90% identity DRGV_70_proteins.tar.gz : protein families at 70% identity DRGV_50_proteins.tar.gz : protein families at 50% identity DRGV_30_proteins.tar.gz : protein families at 30% identity Each archive contains, with {level} being unique, 100, 95, 90, 70, 50 or 30: DRGV_{level}_proteins.rep_seq.faa : representative amino-acid sequence of each family DRGV_{level}_proteins.cluster_info.tsv : family membership at this level DRGV_{level}_proteins.cluster_info_cumulative.tsv : family membership traced back to the individual proteins Three uncompressed tables accompany them: DRGV_protein_header_map.tsv : maps every protein prefix to its contig, genome and vOTU (85,619 rows) DRGV_phold_per_cds_predictions.tsv : PHOLD annotation per CDS - coordinates, strand, PHROG group, function, product and annotation method, with the vOTU, genome and contig it belongs to (1,480,912 rows) DRGV_phold_all_cds_functions.tsv : PHOLD function counts per contig, with the same vOTU and genome identifiers (1,369,872 rows)



