Bacillus pseudomycoides CHAES I 2_2 Prokka genome annotation
收藏资源简介:
A taxonomically targeted reference dataset was constructed for members of the family Bacillaceae. Taxonomic identifiers (taxIDs) corresponding to Bacillaceae species were obtained using the ETE3 NCBI Taxonomy toolkit, which locally mirrors the NCBI taxonomy database. The taxonomy database was updated using NCBITaxa().update_taxonomy_database(), after which all descendant taxIDs belonging to the Bacillaceae family were retrieved programmatically and saved as a plain-text list.Based on this taxID list, publicly available genome assemblies were downloaded from the NCBI GenBank database using the ncbi-genome-download utility. Genome assemblies classified as complete genome, chromosome, or scaffold level were retained. Downloaded data were organized in the standard GenBank directory structure, where each assembly directory contained genomic FASTA files (*_genomic.fna.gz), annotated protein sequences (*_protein.faa.gz), and corresponding GenBank annotation files (*_genomic.gbff.gz). For downstream functional annotation, annotated protein sequences from the downloaded Bacillaceae genomes were used as a custom reference dataset. Protein FASTA files (*_protein.faa.gz) were extracted from all GenBank assembly directories and combined into a single reference protein collection. This dataset served as the taxonomically restricted protein database for homology-based annotation.Genome annotation was performed using Prokka v1.15.6, executed within a Docker container (staphb/prokka:latest) to ensure software reproducibility and to avoid dependency conflicts associated with local installations.The target genome assembly (CHAES_2_2.fna) was annotated using Prokka in bacterial annotation mode. The annotation pipeline included:Prodigal for coding sequence predictionBLASTP searches against the custom Bacillaceae protein datasetIntegration of genus-level annotation heuristics where applicableAnnotation was executed with multi-threading enabled to improve performance. Output files included annotated GFF3, GenBank, protein FASTA, and nucleotide FASTA files, generated in a dedicated output directory.
本研究针对芽孢杆菌科(Bacillaceae)类群构建了分类学靶向参考数据集。我们借助本地镜像NCBI分类学数据库的ETE3 NCBI分类学工具包,获取了芽孢杆菌科物种对应的分类学标识符(taxIDs)。首先使用NCBITaxa().update_taxonomy_database()方法更新分类学数据库,随后以编程方式检索并提取芽孢杆菌科的所有下级分类学标识符,将其保存为纯文本列表。基于该分类学标识符列表,我们使用ncbi-genome-download工具从NCBI GenBank数据库中下载公开可用的基因组组装数据,仅保留标注为完成基因组、染色体或支架序列(scaffold)水平的基因组组装结果。下载的数据按照标准GenBank目录结构进行组织,每个组装目录均包含基因组FASTA文件(*_genomic.fna.gz)、注释后的蛋白质序列文件(*_protein.faa.gz)以及对应的GenBank注释文件(*_genomic.gbff.gz)。为开展下游功能注释,我们将下载的芽孢杆菌科基因组的注释蛋白质序列作为自定义参考数据集:从所有GenBank组装目录中提取蛋白质FASTA文件,并合并为一个统一的参考蛋白质集,该数据集可作为分类学范围受限的蛋白质数据库,用于基于同源性的注释工作。基因组注释使用Prokka v1.15.6完成,运行于Docker容器(staphb/prokka:latest)中,以确保软件可复现性并避免本地安装带来的依赖冲突。我们使用Prokka的细菌注释模式对目标基因组组装结果CHAES_2_2.fna进行注释,该注释流程包含:用于编码序列预测的Prodigal、针对自定义芽孢杆菌科蛋白质数据集的BLASTP搜索,以及适用时整合属级注释启发式规则。注释运行时启用多线程以提升性能,输出文件包含注释后的GFF3、GenBank、蛋白质FASTA以及核苷酸FASTA文件,生成于专属的输出目录中。



