遇见数据集

FMDB, QSFMB and QQFMB databases

收藏
Zenodo2024-03-23 更新2026-05-29 收录
官方服务:

资源简介:

DOI: 10.5281/zenodo.10806681 A catalogue of 10.5M proteins from bacterial species present in food microbiomes was constructed and used to refine and construct biome-specific databases of quorum sensing and quorum quenching related genes. Methods Identification of bacterial species present in food microbiomes The FoodMicrobionet collates data from metataxonomic studies of over 10,000 food and environmental samples (Parente, Zotta & Ricciardi, 2022) and was used to select 4,507 taxa identified to species level (taxa_list.txt). Taxonkit (version 0.5.0) was used to infer taxonomy IDs (NCBI:txid) of 4,264 taxa (this could not be determined for 245 taxa, most frequently for groups of uncertain taxonomic standing and Candidatus species, "taxa_taxon.txt"). All available reference genomes (n=3,478) were downloaded from NCBI (January 2023) via the ncbi-datasets CLI, removing homotypic synonyms due to taxonomic revision where more than one species mapped to the same TaxID ("removed_synonyms.txt"). Generation of protein catalogue of Food Microbiomes (FMDB) Prodigal (version 2.6.3) was used to predict the proteome of each reference genome, which was concatenated together. This pan-proteome was clustered at a sequence identity threshold of 0.9 using CD-HIT (version 4.7), yielding a total of 10,575,338 predicted proteins. Quorum Sensing in Food Microbiomes database (QSFM) Table D1 from the Quorum Sensing of Human Gut Microbiomes (QSHGM; Wu et al, 2022) was used as a template of validated quorum sensing-related genes and expanded by incorporating recently elucidated genes to a total of 240, including gene name; UniProtID; Type of Quorum sensing system (4 levels); whether it was a synthase or receptor. The amino acid sequence of each UniProtID was retrieved using the Unipressed client.,concatenated together and headers were adjusted using the AdjustFastaHeadersForShortBRED.py script. Diamond was used to query the previously-generated FMDB using the concatenated sequences ("qs_multi_unique.fa"). Clustering at 0.90 sequence identity threshold produced 2492 sequences, while at 0.50 sequence identity threshold CD-HIT produced a set of 932 proteins. Quorum Quenching in Food Microbiomes database (QQFM) Table S1 from a recent review by Sikdar & Elias (2020) was manually curated to retrieve, where available, UniprotIDs for validated quorum quenching proteins (n=82). Amino acid sequences were retrieved for each ID using the Unipressed client, concatenated, and headers were adjusted using the AdjustFastaHeadersForShortBRED.py script. Diamond was used to query the previously-generated FMDB using the concatenated sequences ("qq_multi_unique.fa"). Clustering at 0.90 sequence identity threshold produced 1268 sequences, while at 0.50 sequence identity threshold CD-HIT produced a set of 344 proteins. References Parente E, Zotta T, Ricciardi A. FoodMicrobionet v4: A large, integrated, open and transparent database for food bacterial communities. Int J Food Microbiol. 2022 Jul 2;372:109696. doi: 10.1016/j.ijfoodmicro.2022.109696. Epub 2022 May 2. PMID: 35526357. Sikdar R, Elias M. Quorum quenching enzymes and their effects on virulence, biofilm, and microbiomes: a review of recent advances. Expert Rev Anti Infect Ther. 2020 Dec;18(12):1221-1233. doi: 10.1080/14787210.2020.1794815. Epub 2020 Aug 4. PMID: 32749905; PMCID: PMC7705441. Wu S, Feng J, Liu C, Wu H, Qiu Z, Ge J, Sun S, Hong X, Li Y, Wang X, Yang A, Guo F, Qiao J. Machine learning aided construction of the quorum sensing communication network for human gut microbiota. Nat Commun. 2022 Jun 2;13(1):3079. doi: 10.1038/s41467-022-30741-6. PMID: 35654892; PMCID: PMC9163137.

DOI: 10.5281/zenodo.10806681 本研究构建了食品微生物组中细菌物种的1050万个蛋白质目录,并将其用于优化并构建群体感应(quorum sensing, QS)与群体淬灭(quorum quenching, QQ)相关基因的菌群特异性数据库。 ## 研究方法 ### 食品微生物组中细菌物种的鉴定 FoodMicrobionet整合了超过10000份食品及环境样本的宏分类学(metataxonomic)研究数据(Parente、Zotta与Ricciardi,2022),并用于筛选出4507个鉴定至物种水平的分类单元("taxa_list.txt")。使用Taxonkit(版本0.5.0)为4264个分类单元推断分类学ID(NCBI:txid);245个分类单元无法确定其ID,此类情况多出现于分类学地位不明的类群及候选(Candidatus)物种,相关信息存储于"taxa_taxon.txt"。通过ncbi-datasets命令行界面于2023年1月从NCBI下载所有可用的参考基因组(共3478个),并移除因分类学修订产生的同型同义词——即多个物种映射至同一TaxID的情况,移除的同义词列表存储于"removed_synonyms.txt"。 ### 食品微生物组蛋白质目录(Food Microbiomes Database, FMDB)的构建 使用Prodigal(版本2.6.3)预测每个参考基因组的蛋白质组,并将所有预测的蛋白质组进行拼接。使用CD-HIT(版本4.7)以90%的序列同一性阈值对该泛蛋白质组(pan-proteome)进行聚类,最终得到总计10575338个预测蛋白质。 ### 食品微生物组群体感应数据库(Quorum Sensing in Food Microbiomes, QSFM) 以《人类肠道微生物组群体感应》(Quorum Sensing of Human Gut Microbiomes, QSHGM; Wu等,2022)中的表D1作为已验证群体感应相关基因的模板,并通过纳入新近阐明的基因将其扩充至240个,包含基因名称、UniProtID、群体感应系统类型(4个层级)以及该基因是否为合酶或受体。使用Unipressed客户端检索每个UniProtID对应的氨基酸序列,将序列拼接后通过AdjustFastaHeadersForShortBRED.py脚本调整序列标题。使用Diamond以拼接后的序列("qs_multi_unique.fa")对先前构建的FMDB进行序列比对检索。以90%序列同一性阈值聚类得到2492条序列,以50%序列同一性阈值通过CD-HIT聚类得到932个蛋白质。 ### 食品微生物组群体淬灭数据库(Quorum Quenching in Food Microbiomes, QQFM) 手动整理Sikdar与Elias(2020)近期综述中的表S1,检索其中已验证的群体淬灭蛋白质的UniProtID(共82个,若存在对应ID)。使用Unipressed客户端检索每个ID对应的氨基酸序列,将序列拼接后通过AdjustFastaHeadersForShortBRED.py脚本调整序列标题。使用Diamond以拼接后的序列("qq_multi_unique.fa")对先前构建的FMDB进行序列比对检索。以90%序列同一性阈值聚类得到1268条序列,以50%序列同一性阈值通过CD-HIT聚类得到344个蛋白质。 ## 参考文献 1. Parente E, Zotta T, Ricciardi A. FoodMicrobionet v4: 一款面向食品细菌群落的大型整合型开放透明数据库. 国际食品微生物学杂志, 2022 Jul 2;372:109696. DOI: 10.1016/j.ijfoodmicro.2022.109696. 2022年5月2日在线发表. PMID: 35526357. 2. Sikdar R, Elias M. 群体淬灭酶及其对毒力、生物被膜与微生物组的影响:最新进展综述. 专家评论抗感染治疗, 2020 Dec;18(12):1221-1233. DOI: 10.1080/14787210.2020.1794815. 2020年8月4日在线发表. PMID: 32749905; PMCID: PMC7705441. 3. Wu S, Feng J, Liu C, Wu H, Qiu Z, Ge J, Sun S, Hong X, Li Y, Wang X, Yang A, Guo F, Qiao J. 机器学习辅助构建人类肠道菌群群体感应通信网络. 自然通讯, 2022 Jun 2;13(1):3079. DOI: 10.1038/s41467-022-30741-6. PMID: 35654892; PMCID: PMC9163137.

提供机构:
Zenodo
创建时间:
2024-03-23
二维码
社区交流群
二维码
科研交流群
商业服务