Tree of life (archaea and bacteria only) and MAGs isolated from two hot springs in the Uzon Caldera, Kamchatka, Russia (tre file)
收藏资源简介:
Taxonomy of our MAGs was refined by placing MAGs in a phylogenetic context using PhyloSift v. 1.0.1 with the updated PhyloSift markers database (version 4; 2018-02-12; https://figshare.com/articles/PhyloSift_markers_database/5755404/4). For this purpose, MAGs, all taxa previously identified by Burgess et al. (2012) with complete genomes available on NCBI (downloaded 2017-09-06) and all archaeal and bacterial genomes previously used in Hug et al. (2016) were placed in a phylogenetic tree. All genomes used in this tree and a mapping file can be found on figshare (https://doi.org/10.6084/m9.figshare.6863594.v1; https://doi.org/10.6084/m9.figshare.6863744.v2; https://doi.org/10.6084/m9.figshare.6863798.v1; https://doi.org/10.6084/m9.figshare.6863813.v1). For more information about the available core marker gene sets see Darling et al. 2014 (updated last on 2018-02-12). We used 37 of these single-copy marker genes (ribosomal protein S2 rpsB, S10 rpsJ, L1 rplA, L22, L4/L1e rplD, L2 rplB, S9 rpsl, L3 rplC, L14b/L23e rplN, S5, S19 rpsS, S7, L16/L10E rplP, S13 rpsM, L15, L25/L23, L6 rplF, L11 rplK, L5 rplE, S12/S23, L29, S3 rpsC, S11 rpsK, L10, S8, L18P/L5E, S15P/S13e, S17, S13 rplM, L24; and translation initiation factor IF-2, metalloendopeptidase, phenylalanyl-tRNA synthetase beta subunit, phenylalanyl-tRNA synthetase alpha subunit, tRNA pseudouridine synthase B, Porphobilinogen deaminase, and ribonuclease HII; i.e., PhyloSift markers DNGNGWU00001 - DNGNGWU00040 without DNGNGWU00004, DNGNGWU00008 and DNGNGWU00038). The amino acid alignment of these 37 concatenated genes was trimmed using trimAl v.1.2. Columns with gaps in more than 5% of the sequences were removed, as well as taxa with with less than 75% of the concatenated sequences. MAGs from ARK and ZAV that did not meet this threshold were manually kept in the alignment. The final alignment comprised 3,240 taxa (make a supplementary table) and 5,459 amino acid positions. This alignment was then used to build a new phylogenetic tree in RAxML v. 8.2.10 on the CIPRES Science Gateway web server. First, we searched for the best protein substitution model of the alignment with its empirical base frequencies using the Bayesian Information Criterion (BIC) within RAxML. Then, this substitution model; i.e, LG plus CAT (after Le and Gascuel), was used to infer a phylogenetic tree. We chose the rapid bootstrapping algorithm (flag -f a) to find the best scoring maximum likelihood tree with 10 starting trees in one run with the number of bootstraps automatically determined (MRE-based bootstopping criterion). One hundred fifty bootstrap replicates were conducted. The full tree inference required 2,236 computational hours on the CIPRES supercomputer.
本研究通过将宏基因组组装基因组(Metagenome-Assembled Genomes, MAGs)置于系统发育背景下,对其分类学信息进行精细化修正:采用PhyloSift v.1.0.1及更新版PhyloSift标记基因数据库(版本4;2018-02-12;https://figshare.com/articles/PhyloSift_markers_database/5755404/4)完成分析。本次分析纳入的序列包括本研究的MAGs、Burgess等人(2012)已鉴定且在NCBI数据库中可获取完整基因组(2017-09-06下载)的所有类群,以及Hug等人(2016)研究中使用的全部古菌与细菌基因组,并将上述序列整合至系统发育树构建流程中。本研究使用的全部基因组及映射文件均可从figshare平台获取(https://doi.org/10.6084/m9.figshare.6863594.v1;https://doi.org/10.6084/m9.figshare.6863744.v2;https://doi.org/10.6084/m9.figshare.6863798.v1;https://doi.org/10.6084/m9.figshare.6863813.v1)。如需了解可用核心标记基因集的详细信息,可参考Darling等人(2014)的研究(最后更新于2018-02-12)。 本研究选取上述数据库中的37个单拷贝标记基因,具体包括:核糖体蛋白S2(rpsB)、S10(rpsJ)、L1(rplA)、L22、L4/L1e(rplD)、L2(rplB)、S9(rpsl)、L3(rplC)、L14b/L23e(rplN)、S5、S19(rpsS)、S7、L16/L10E(rplP)、S13(rpsM)、L15、L25/L23、L6(rplF)、L11(rplK)、L5(rplE)、S12/S23、L29、S3(rpsC)、S11(rpsK)、L10、S8、L18P/L5E、S15P/S13e、S17、S13(rplM)、L24;以及翻译起始因子IF-2、金属内肽酶、苯丙氨酰-tRNA合成酶β亚基、苯丙氨酰-tRNA合成酶α亚基、tRNA假尿苷合酶B、胆色素原脱氨酶与核糖核酸酶HII;即PhyloSift标记基因DNGNGWU00001至DNGNGWU00040,排除DNGNGWU00004、DNGNGWU00008与DNGNGWU00038。 使用trimAl v.1.2对这37个基因的串联氨基酸序列比对结果进行修剪:移除在超过5%的序列中存在缺失位点的列,同时移除串联序列覆盖度低于75%的类群。但来自ARK与ZAV的未达上述阈值的MAGs被手动保留于比对结果中。最终的比对结果包含3240个类群(详见补充表格)与5459个氨基酸位点。 随后,以该比对结果为基础,在CIPRES科学网关(CIPRES Science Gateway)服务器上通过RAxML v.8.2.10构建全新的系统发育树。首先,在RAxML中基于贝叶斯信息准则(Bayesian Information Criterion, BIC)结合经验碱基频率,为该序列比对筛选最优的蛋白质替换模型;最终选定LG+CAT替换模型(源自Le与Gascuel的研究)用于系统发育树推断。本研究采用快速自举算法(参数-f a),单次运行中使用10个初始树,并基于MRE自举停止准则自动确定自举重复次数,最终完成150次自举重复。本次完整的系统发育树推断过程在CIPRES超级计算机上总计耗时2236个计算小时。



