Genome-scale community modelling reveals key metabolic cross-feedings in epipelagic bacterioplankton communities (Supplementary Materials)
收藏资源简介:
A comprehensive catalog of 19,791 marine prokaryotic isolates (WGS), single-amplified genomes (SAGs) and metagenomic-assembled genomes (MAGs) compiled from MarRef v4.0 (N=943, mostly high-quality WGS), MarDB v4.0 (N=12,963), and the aquatic representative genomes from the ProGenomes database v1.0 (N=566). This collection of well-documented genomes was complemented by 5,319 MAGs assembled from four distinct studies, namely: Parks et al. 2017 (DOI; N=1,765; downloaded from EBI), Tully et al. 2017/2018 (DOI and DOI; N=2,597; downloaded from EBI), and Delmont et al. 2018 (DOI; N=957; downloaded from FIGSHARE). The Parks et al. study contained genomes reconstructed from non-marine biomes. Thus, a selection of 1,765 genomes was extracted by searching for specific keywords: “tara|marine|sea|ocean|mediterranean” (case insensitive). Note that depending on their study of origin, included MAGs may have been reconstructed using different assembling and binning methods. The archive includes: a metadata file describing the quality and redundancy of the genomes named `EcoSysMic_metadata.tsv` sequences of the 19,791 (redundant) genomes in `All/WGS` companion files in `All/Data` and `dRep95/Data` (see Methods in the associated paper), including predicted CDS and EggNOG functional annotations predicted GTDB taxonomy CarveMe reconstructed metabolic models and their MEMOTE quality The 7,658 non-redundant species-level genomes (delineated by a 95% ANI threshold over 60% of genome length) that were used in the associated paper are defined by the column `is_drep95` in the metadata file.
本数据集为一套涵盖19791株海洋原核生物分离株全基因组测序(Whole Genome Sequencing, WGS)、单扩增基因组(Single-Amplified Genomes, SAGs)与宏基因组组装基因组(Metagenome-Assembled Genomes, MAGs)的综合目录,数据整合自三个公开来源:MarRef v4.0(样本量N=943,以高质量WGS为主)、MarDB v4.0(N=12963),以及ProGenomes数据库v1.0中的水生代表性基因组(N=566)。这套经过详实记录的基因组集合,额外补充了来自四项独立研究的5319个MAGs,具体包括:Parks等人2017年的研究(DOI;N=1765,从EBI平台下载)、Tully等人2017/2018年的两项相关研究(两篇DOI;N=2597,从EBI平台下载),以及Delmont等人2018年的研究(DOI;N=957,从FIGSHARE平台下载)。由于Parks等人的研究包含了非海洋生物群落中重建的基因组,因此本数据集通过检索不区分大小写的关键词“tara|marine|sea|ocean|mediterranean”,从其原始的1765个基因组中筛选得到符合海洋来源的子集。需注意的是,根据来源研究的不同,本次纳入的MAGs可能采用了各异的基因组组装与分箱方法。本归档文件包含以下内容: 1. 一份名为`EcoSysMic_metadata.tsv`的元数据文件,用于描述各基因组的质量与冗余度; 2. 19791个(含冗余)基因组的序列数据,存储于`All/WGS`、`All/Data`及`dRep95/Data`配套文件中(详细说明参见相关论文的方法部分),其中包含预测的编码序列(Coding Sequence, CDS)、EggNOG功能注释信息,以及基于GTDB(Genome Taxonomy Database)的分类学预测注释; 3. CarveMe工具重建的代谢模型及其MEMOTE质量评估结果。本相关论文中使用的7658个非冗余物种级基因组(以95%平均核苷酸一致性(Average Nucleotide Identity, ANI)阈值、覆盖60%基因组长度的标准划定),可通过元数据文件中的`is_drep95`列进行快速定位筛选。



