Microbial genome collection of aerobic granular sludge cultivated in sequential batch reactor using different carbon source mixtures
收藏资源简介:
Aerobic granular sludge microogranisms are cultivated in a sequential batch reactor (SRB), and a metagenome-assembled genome has been successfully acquired through the use of Illumina short-read and PacBio long-read metagenomics. This genome collection builds on a prior study (PRJEB38840; Aline adler, Sfam 2022) by obtaining raw sequences of short-read and long-read metagenomics. This study utilizes the dataset to construct a hybrid assembly of Illumina short reads (two samples: DNA extraction A and B) and PacBio long reads (one sample: DNA extraction A) using the SPAdes assembler. The resulting metagenomic assemblies are presented for each of the days (71, 322, 427, 740), and samples were fed different carbon source mixtures, including volatile fatty acids (day 71), complex monomeric (day 322 and 427), and complex polymeric (day 740). The study also conducts a comparative analysis of the new and old MAGs and provides the combined dataset to the public (n=759). These MAGs have been dereplicated to obtain representatives (n=233). The whole MAG collection is indexed using the gtdb taxonomy and are accompanied by key metrics such as genome size, number of contigs, and N50 value. Additionally, the study offers R-codes for data analysis.
本数据集以序批式反应器(Sequential Batch Reactor, SRB)中培养的好氧颗粒污泥微生物为研究对象,通过Illumina短读长测序与PacBio长读长宏基因组测序技术,成功获取了宏基因组组装基因组(Metagenome-Assembled Genome, MAG)。本次基因组集合基于此前一项研究(PRJEB38840;Aline Adler, Sfam 2022),补充获取了短读长与长读长宏基因组测序的原始序列数据。本研究依托该数据集,使用SPAdes组装工具,对Illumina短读长数据(2个样本:DNA提取批次A与B)与PacBio长读长数据(1个样本:DNA提取批次A)进行混合组装。针对培养天数71、322、427、740的样本分别提供了最终得到的宏基因组组装结果;各阶段样本的投喂碳源组合分别为:第71天为挥发性脂肪酸(Volatile Fatty Acids, VFA),第322、427天为复合单体碳源,第740天为复合聚合碳源。本研究同时对新旧宏基因组组装基因组进行了比较分析,并将合并后的公开数据集(n=759)进行了共享。这些宏基因组组装基因组已经过去冗余处理,得到了233个代表基因组(n=233)。整套宏基因组组装基因组集合已基于基因组分类数据库(Genome Taxonomy Database, GTDB)的分类学进行了索引,并附带了基因组大小、重叠群数量、N50值等关键统计指标。此外,本研究还提供了用于数据分析的R语言代码脚本。



