遇见数据集

Complete gene content and sequences for all Hpar samples in CHUVI pangenomic analysis study

收藏
Zenodo2025-08-19 更新2026-05-29 收录
官方服务:

资源简介:

259 genomes of Hamophilus parainfluezae were annotated, with genes assigned to a KEGG pathway. The following dataset includes, for each sample, all genes with their protein sequence, gene description and KEGG pathway assigned. This output was generated using prokka (for genetic annotation and protein sequence generation) and KEGGREST (for KEGG ortholog and pathway assignment). Columns of the dataset work as follows: SampleID: Sample ID from the SRA database (NCBI) Columns from prokka output: locus_tag: unique identifier for each sequence generated by prokka ftype: type of genomic feature (CDS = coding sequence; CRISPR; tRNA = transference RNA; tmRNA transfer-messenger RNA) length_bp: length of genomic feature in base pairs gene: gene name EC_number: Enzyme Comission number COG: Cluster of Ortholog Groups family Sequence: protein aminoacid sequence (for CDS types) Columns from KEGGREST output: ortholog: KEGG Orthology database identifier Pathway: KEGG Pathway database identifier

提供机构:
Zenodo
创建时间:
2025-06-23
二维码
社区交流群
二维码
科研交流群
商业服务