Fermented Foods Microbial Genomes Database
收藏资源简介:
Fermented Foods Microbial Genomes Database This database contains 13,850 microbial genomes assembled from various fermented foods and associated curated metadata. We have also clustered the database at 95% identity for creating a species-representative database, and at 99% identity for creating a "strain"-representative database, since we hypothesize that many bioactivities and phenotypes for fermented food microbes are important at the strain-level. This GitHub repository documents how publicly available genomes and metagenome-assembled genomes were sourced and curated. This GitHub repository documents how the associated metadata was curated. This database largely pulls from existing genome resources, and we curated this database specifically for fermented foods. If you use this database, please cite the following genome databases/resources: MiFoDB, a workflow for microbial food metagenomic characterization, enables high-resolution analysis of fermented food microbial dynamics. Elisa B. Caffrey, Matthew R. Olm, Caroline Isabel Kothe, Joshua Evans, Justin L. Sonnenburg . bioRxiv 2024.03.29.587370; doi: https://doi.org/10.1101/2024.03.29.587370 Unexplored microbial diversity from 2,500 food metagenomes and links with the human microbiome. Carlino, Niccolo Alvarez-Ordonez, Avelino et al. Cell, Volume 187, Issue 20, 5775-5795.e15 AND the associated Zenodo release: Master Consortium. (2024). Unexplored microbial diversity from 2,500 food metagenomes and links with the human microbiome. Zenodo. https://doi.org/10.5281/zenodo.13285428 We are incredibly grateful for these groups and countless others taking the time to make their data publicly available. Included in the metadata is the original DOI and study link from which the genome was generated, in addition to if they were collated into one of the above two larger databases. If you specifically use/analyze a subset of genomes, please cite those studies to credit those that generate data and make it publicly available. Subsetting the Database to Species/Strain-Resolved Representatives or a Custom Set We have provided the entire set of 13,850 microbial genomes in a single tar archive for download. We have also provided tar archives of the genomes clustered at 95% and 99% identity. If you wish to download the entire database and then only use a subset of the database, such as species-representative (clustered at 95% ANI) or "strain-representative" (clustered at 99% ANI) genomes after downloading the entire database, you can use our helper script for subsetting genomes that are representatives or a custom list that you provide. usage: subset_genomes.py [-h] [--rep-column {rep_95id,rep_99id}] [--id-column ID_COLUMN] [--dry-run] [--genome-list GENOME_LIST] metadata_tsv all_genomes_dir output_dir Subset representative genomes (species/strain) from a genome set usingmetadata. positional arguments: metadata_tsv Path to metadata TSV file all_genomes_dir Directory containing all .fa genome files output_dir Directory to copy representative genomes to optional arguments: -h, --help show this help message and exit --rep-column {rep_95id,rep_99id} Column in metadata to use for representatives (e.g., rep_95id or rep_99id) --id-column ID_COLUMN Column in metadata with genome file IDs (default: mag_id) --dry-run Only print what would be copied, don't actually copy --genome-list GENOME_LIST Optional: Path to file with list of genome IDs or filenames to subset (one per line) KBase Fermented Foods Microbial Genomes Database Narrative We have uploaded the "strain-representative" set of ~4300 genomes to KBase as a public narrative. KBase is a community-driven platform for facilitating open science research in systems biology. KBase allows you to run bioinformatics tools on large datasets using freely available Department of Energy high-perofrmance computing resources, allowing for open-sharing of research outputs and collaborative work. You can sign-up for a KBase account with any email account. You are not required to be affiliated with an academic institution or government organization to use KBase, and you can reside outside of the United States. This platform not only allows additional access to the Fermented Foods Microbial Genomes Database, but access to open-source bioinformatics tools and high-performance computing resources through the DOE to run reproducible analyses. You can use this narrative for exploring the database, incorporating your own genomes to compare against genomes in the database, and/or using as a teaching resource.
# 发酵食品微生物基因组数据库(Fermented Foods Microbial Genomes Database) 本数据库包含13850条从各类发酵食品中组装得到的微生物基因组,以及配套的经过整理注释的元数据。 我们还将该数据库按95%序列一致性进行聚类,以构建物种代表性数据库;按99%序列一致性进行聚类,以构建“菌株”代表性数据库——这一设计基于我们的假设:发酵食品微生物的诸多生物活性与表型特征在菌株层面具有关键意义。 本GitHub仓库详细记录了公开可用基因组及宏基因组组装基因组(metagenome-assembled genomes, MAGs)的获取与整理流程,同时也记录了配套元数据的整理方法。 本数据库主要依托现有基因组资源构建,我们专门针对发酵食品领域对其进行了系统性整理。若您使用本数据库,请引用以下基因组数据库/资源: 1. MiFoDB:一款用于微生物食品宏基因组表征的分析流程,可实现发酵食品微生物群落动态的高分辨率解析。作者:Elisa B. Caffrey、Matthew R. Olm、Caroline Isabel Kothe、Joshua Evans、Justin L. Sonnenburg。预印本发布于bioRxiv,2024年3月29日,编号2024.03.29.587370;DOI:https://doi.org/10.1101/2024.03.29.587370 2. 《来自2500份食品宏基因组的未被探索的微生物多样性及其与人类微生物组的关联》,作者:Carlino、Niccolo Alvarez-Ordonez、Avelino等。发表于《Cell》,第187卷第20期,页码范围5775-5795.e15;以及配套的Zenodo开源资源:Master Consortium. (2024). Unexplored microbial diversity from 2,500 food metagenomes and links with the human microbiome. Zenodo. https://doi.org/10.5281/zenodo.13285428 我们衷心感谢上述团队以及无数其他愿意将研究数据公开共享的科研人员。元数据中包含了生成对应基因组的原始DOI及研究链接,同时标注了该基因组是否整合自上述两个大型数据库之一。若您仅使用或分析部分基因组子集,请引用对应原始研究文献,以表彰生成并公开这些数据的研究者。 ## 数据库子集提取:物种/菌株分辨率代表性基因组或自定义基因组集 我们提供了包含全部13850条微生物基因组的单个tar归档文件供下载,同时也提供了按95%和99%序列一致性聚类后的基因组tar归档文件。若您希望先下载完整数据库,再仅使用其中子集(例如按95%平均核苷酸一致性(Average Nucleotide Identity, ANI)聚类得到的物种代表性基因组,或按99% ANI聚类得到的“菌株代表性”基因组),可使用我们提供的辅助脚本完成代表性基因组子集提取,或基于您提供的自定义列表进行精准子集提取。 ### subset_genomes.py 脚本使用说明 #### 用法 subset_genomes.py [-h] [--rep-column {rep_95id,rep_99id}] [--id-column ID_COLUMN] [--dry-run] [--genome-list GENOME_LIST] metadata_tsv all_genomes_dir output_dir #### 功能 基于元数据从基因组集中提取代表性基因组(物种/菌株级别) #### 参数说明 ##### 位置参数 metadata_tsv 元数据TSV(Tab-Separated Values,制表符分隔值)格式文件的存储路径 all_genomes_dir 包含所有.fa格式基因组文件的目录路径 output_dir 用于复制目标代表性基因组的输出目录路径 ##### 可选参数 -h, --help 显示帮助信息并退出脚本 --rep-column {rep_95id,rep_99id} 元数据中用于标识代表性基因组的列名(例如rep_95id或rep_99id) --id-column ID_COLUMN 元数据中包含基因组文件ID的列名(默认值:mag_id) --dry-run 仅打印将要执行的复制操作列表,不实际执行文件复制 --genome-list GENOME_LIST 可选参数:包含待提取基因组ID或文件名列表的文件路径(每行一个条目) ## KBase发酵食品微生物基因组数据库叙事页面(KBase Fermented Foods Microbial Genomes Database Narrative) 我们已将约4300条“菌株代表性”基因组集上传至KBase平台,作为公开叙事页面。 KBase是一个由社区驱动的开源平台,旨在推动系统生物学领域的开放科学研究。该平台允许用户利用美国能源部提供的免费高性能计算资源,对大型数据集运行生物信息学分析工具,从而实现研究成果的开放共享与跨团队协作研究。您可使用任意电子邮箱注册KBase账号,无需隶属于学术机构或政府组织,也无需身处美国境内。 该平台不仅为用户提供了访问发酵食品微生物基因组数据库的额外途径,还可通过美国能源部的基础设施访问开源生物信息学工具与高性能计算资源,以开展可重复的科研分析工作。您可使用该叙事页面探索本数据库,将您自主生成的基因组纳入分析以与数据库内的基因组进行比对,或将其用作教学辅助资源。



