Preprint of the Supplementary Material of the manuscript entitled "Comparative genomics of Stutzerimonas balearica (Pseudomonas balearica): diversity, habitats and biodegradation of aromatic compounds"
收藏资源简介:
This dataset contains the Supplementary Material of the manuscript entitled "Comparative genomics of Stutzerimonas balearica (Pseudomonas balearica): diversity, habitats and biodegradation of aromatic compounds", which has not been certified by peer review and which is part of the PhD thesis of Francisco Salvà Serra, entitled "Bacterial whole-genome sequencing for establishment of reference sequences, comparative genomics, biomarker discovery and characterization of novel taxa". The PhD thesis is part of the PhD programme in Environmental and Biomedical Microbiology of the University of the Balearic Islands (Spain). Description of the files: Supplementary Figure 1. Phylogenetic tree based on partial rpoD sequences (549 bp). The tree includes 111 of the 113 genome sequences included in the study; the other two did not contain any rpoD sequence in the assembly. The phylogenetic analysis also includes the reference rpoD sequences of most of the described species and phylogenomic species of Stutzerimonas. The GenBank accession numbers of the genome sequences are listed in Supplementary Table 1 and those of the reference sequences obtained from the PseudoMLSA database are indicated in parenthesis. For each strain, the final taxonomic assignment is indicated. The final phylogenomic species assignment of each strain is indicated in parenthesis. The distances were calculated, using the Jukes-Cantor method, and the tree was constructed, using the neighbor-joining method. Bootstrap values of 50% or greater (from 1,000 replicates) are shown at the nodes. P. aeruginosa CCM 1960T was used as an outgroup. Pgs: phylogenomic species; gv.: genomovar; ref.: reference. Supplementary Figure 2. Core genome-based phylogenomic tree of the 11 isolate-derived genome sequences of S. balearica, 81 strains listed as S. stutzeri in NCBI and the type strains of 11 additional species of the genus Stutzerimonas. The tree was generated, based on 210,488 homologous amino acid positions, derived from 699 single copy core proteins. Pgs: phylogenomic species; ref.: reference. Supplementary Table 1. List of the 113 genome sequences used in this study and their associated metadata. Supplementary Table 2. ANIb values between the 113 genome sequences (all vs. all) included in this study. Each pairwise comparison was performed bidirectionally; the average value is displayed. Supplementary Table 3. List of the 176 confirmed S. balearica strains and their associated metadata. Supplementary File 1. Metagenome assembled genome (MAG) sequence of S. balearica UBA3230, annotated using DFAST, in GBK format. Supplementary File 2. Metagenome assembled genome (MAG) sequence of S. balearica UBA6635, annotated using DFAST, in GBK format. Supplementary File 3. Metagenome assembled genome (MAG) sequence of S. balearica 3300027365_7, annotated using DFAST, in GBK format. Supplementary File 4. Genomic features of the genome sequences of S. balearica. Sheet 1: Pan-genes of S. balearica. Sheet 2: Genes encoding enzymes for catabolism of aromatic compounds detected in one or more genome sequences of S. balearica. Sheet 3: Genes for catabolism of aromatic compounds not detected in the genome sequences of S. balearica. Sheet 4: Antibiotic resistance genes detected, using Resistance Gene Identifier (RGI) and the Comprehensive Antibiotic Resistance Database (CARD). Sheet 5: Biocide and metal resistance genes predicted by BacMet. Sheet 6: Virulence determinants detected, using VFanalyzer. Sheet 7: CRISPR-Cas systems predicted by CRISPRone. Sheet 8: Prophages predicted by PHASTER. Sheet 9: Integrative and conjugative elements (ICEs) and integrative and mobilizable elements (IMEs) predicted by ICEfinder. Sheet 10: genes associated with natural transformation capacity. Sheet 11: Regulatory genes predicted, using P2RP.
本数据集包含题为《巴利亚利卡施特策氏菌(原巴利亚利卡假单胞菌)比较基因组学:芳香族化合物的多样性、生境与生物降解性》的手稿的补充材料。该手稿尚未经过同行评议,属于Francisco Salvà Serra的博士学位论文《用于建立参考序列、比较基因组学、生物标志物发现及新类群表征的细菌全基因组测序》的一部分。该博士学位论文隶属于西班牙巴利阿里群岛大学环境与生物医学微生物学博士培养项目。 ### 文件说明 1. 补充图1:基于部分rpoD基因序列(549 bp)构建的系统发育树。本树纳入了研究中113个基因组序列里的111个,其余2个基因组组装结果中未携带任何rpoD序列。本次系统发育分析同时包含了施特策氏菌属多数已描述物种及系统发育组的参考rpoD基因序列。基因组序列的GenBank登录号详见补充表1,从PseudoMLSA数据库获取的参考序列登录号以括号标注。每株菌株均标注了最终分类学鉴定结果,其系统发育组最终鉴定结果以括号标注。进化距离采用Jukes-Cantor模型计算,系统发育树通过邻接法构建。节点处标注了≥50%的自展值(基于1000次重复抽样)。以铜绿假单胞菌P. aeruginosa CCM 1960^T作为外类群。注:Pgs为系统发育组(phylogenomic species)缩写,gv.为基因组变种(genomovar)缩写,ref.为参考株(reference)缩写。 2. 补充图2:基于核心基因组的系统发育树,包含11株分离自巴利亚利卡施特策氏菌的基因组序列、NCBI数据库中标注为斯图泽假单胞菌(S. stutzeri)的81株菌株,以及施特策氏菌属另外11个物种的模式菌株。本树基于699个单拷贝核心蛋白对应的210488个同源氨基酸位点构建。注:Pgs为系统发育组(phylogenomic species)缩写,ref.为参考株(reference)缩写。 3. 补充表1:本研究使用的113个基因组序列及其关联元数据列表。 4. 补充表2:本研究涵盖的113个基因组序列之间的平均核苷酸一致性(ANIb)值(全对全比对)。每一组两两比对均采用双向比对方式,最终展示为平均后的比对结果。 5. 补充表3:176株经确认的巴利亚利卡施特策氏菌菌株及其关联元数据列表。 6. 补充文件1:巴利亚利卡施特策氏菌UBA3230的宏基因组组装基因组(MAG)序列,采用DFAST完成注释,格式为GBK。 7. 补充文件2:巴利亚利卡施特策氏菌UBA6635的宏基因组组装基因组(MAG)序列,采用DFAST完成注释,格式为GBK。 8. 补充文件3:巴利亚利卡施特策氏菌3300027365_7的宏基因组组装基因组(MAG)序列,采用DFAST完成注释,格式为GBK。 9. 补充文件4:巴利亚利卡施特策氏菌基因组序列的基因组特征信息,包含以下工作表: 1. 巴利亚利卡施特策氏菌的泛基因 2. 在一株或多株巴利亚利卡施特策氏菌基因组序列中检测到的、编码芳香族化合物分解代谢酶的基因 3. 未在巴利亚利卡施特策氏菌基因组序列中检测到的芳香族化合物分解代谢相关基因 4. 采用耐药基因识别器(RGI,Resistance Gene Identifier)及综合抗生素抗性数据库(CARD,Comprehensive Antibiotic Resistance Database)检测到的抗生素抗性基因 5. 采用BacMet预测的生物杀灭剂与金属抗性基因 6. 采用VFanalyzer检测到的毒力决定因子 7. 采用CRISPRone预测的CRISPR-Cas系统 8. 采用PHASTER预测的原噬菌体 9. 采用ICEfinder预测的整合接合元件(ICEs,Integrative and Conjugative Elements)及整合可移动元件(IMEs,Integrative and Mobilizable Elements) 10. 与自然转化能力相关的基因 11. 采用P2RP预测的调控基因



