The use of a metabologenomics approach for the discovery of antimicrobial natural products - Thesis Supplementary Data
收藏资源简介:
This dataset supports a multi-omics investigation into the biosynthetic potential and metabolite production of selected microbial strains under varying culture conditions. The primary hypothesis of this study is that strain repositioning combined with the OSMAC (One Strain Many Compounds) approach and integrated metabolomic and genomic analyses can reveal hidden metabolic diversity in actinobacteria. The dataset contains processed LC-MS/MS metabolomics data, whole genome sequencing (WGS) assemblies, genome annotations, and integrated metabologenomics outputs. Metabolomic data were generated using LC-MS/MS and processed into feature-based molecular networks using GNPS2. These data include MGF files, feature tables, metadata, and annotation outputs from tools such as SIRIUS and CANOPUS. Multivariate and univariate statistical analyses (e.g., PCA, PCoA, random forest classification, ANOVA, PERMANOVA) were performed to identify metabolite patterns associated with specific strains and culture conditions. Genomic data include polished genome assemblies, quality control metrics, and functional annotations, alongside biosynthetic gene cluster (BGC) predictions and comparative genomics analyses. Phylogenomic and taxonomic analyses (ANI, AAI, dDDH) were used to contextualize strain identity and novelty. BiG-SCAPE outputs were used to group BGCs into gene cluster families for comparative analysis. Metabologenomics integration was performed to link metabolomic features with predicted BGCs, enabling the identification of candidate gene-metabolite relationships. These results are provided as putative links which could be prioritised for further drug discovery. The data demonstrate that cultivation conditions significantly influence metabolite production profiles, and that integrating metabolomics with genomics enhances the interpretation of microbial chemical diversity. This dataset can be used for reanalysis of molecular networks, validation of genomic predictions, or further development of metabolite annotation and gene cluster linking approaches. Full methodological details are provided in the associated thesis.
本数据集支持针对不同培养条件下选定微生物菌株的生物合成潜力与代谢产物生成情况的多组学研究。本研究的核心假说为:将菌株重定位策略与OSMAC(One Strain Many Compounds,一菌多产物)方法相结合,并整合代谢组学与基因组学分析,可揭示放线菌中隐藏的代谢多样性。 本数据集包含经过预处理的LC-MS/MS代谢组学数据、全基因组测序(WGS)组装结果、基因组注释信息以及整合的代谢基因组学输出成果。其中,代谢组学数据通过LC-MS/MS技术获取,并借助GNPS2平台处理为基于特征的分子网络,涵盖MGF文件、特征表、元数据以及来自SIRIUS、CANOPUS等工具的注释结果。研究通过多变量与单变量统计分析(如PCA、PCoA、随机森林分类、ANOVA、PERMANOVA),筛选出与特定菌株及培养条件相关的代谢物特征模式。 基因组学数据包含经过抛光打磨的基因组组装结果、质量控制指标、功能注释信息,以及生物合成基因簇(BGC)预测与比较基因组学分析结果。研究采用系统发育基因组学与分类学分析方法(ANI、AAI、dDDH),明确菌株的分类地位与新颖性;借助BiG-SCAPE的输出结果,可将BGC划分为基因簇家族以开展比较分析。通过代谢基因组学整合分析,将代谢组学特征与预测得到的BGC进行关联,从而鉴定出潜在的基因-代谢物关联关系。本结果以推定关联的形式提供,可优先用于后续药物发现研究。 本数据集的相关数据表明,培养条件会显著影响代谢物的生成谱,而整合代谢组学与基因组学分析可提升对微生物化学多样性的解析能力。本数据集可用于分子网络的再分析、基因组预测结果的验证,或是进一步优化代谢物注释与基因簇关联分析方法。完整的实验方法细节可参阅相关学位论文。



