遇见数据集

Alignment for tree of life

收藏
Figshare2018-08-09 更新2026-04-08 收录
官方服务:

资源简介:

This alignment was used to build a tree with our MAGs, all taxa previously identified by Burgess et al. (2012) with complete genomes available on NCBI (downloaded 2017-09-06), and all archaeal and bacterial genomes previously used in Hug et al. (2016). The genomes used in this tree and a mapping file can be found on figshare.<br>(genomes in Hug et al.’s tree of life (2016): https://doi.org/10.6084/m9.figshare.6863594.v1, https://doi.org/10.6084/m9.figshare.6863744.v2, https://doi.org/10.6084/m9.figshare.6863813.v1; genomes from Burgess et al. (2012): https://doi.org/10.6084/m9.figshare.6863798.v1).<br> <br>PhyloSift builds an alignment of the concatenated sequences for a set of core markers for each taxon. We used 37 of these single-copy marker genes (ribosomal protein S2 rpsB, S10 rpsJ, L1 rplA, L22, L4/L1e rplD, L2 rplB, S9 rpsl, L3 rplC, L14b/L23e rplN, S5, S19 rpsS, S7, L16/L10E rplP, S13 rpsM, L15, L25/L23, L6 rplF, L11 rplK, L5 rplE, S12/S23, L29, S3 rpsC, S11 rpsK, L10, S8, L18P/L5E, S15P/S13e, S17, S13 rplM, L24; and translation initiation factor IF-2, metalloendopeptidase, phenylalanyl-tRNA synthetase beta subunit, phenylalanyl-tRNA synthetase alpha subunit, tRNA pseudouridine synthase B, Porphobilinogen deaminase, and ribonuclease HII; i.e., PhyloSift markers DNGNGWU00001 - DNGNGWU00040 without DNGNGWU00004, DNGNGWU00008 and DNGNGWU00038). The amino acid alignment of these 37 concatenated genes was trimmed using trimAl v.1.2. Columns with gaps in more than 5% of the sequences were removed, as well as taxa with with less than 75% of the concatenated sequences. MAGs from ARK and ZAV that did not meet this threshold were manually kept in the alignment.

本次序列比对用于构建包含本研究宏基因组组装基因组(Metagenome-Assembled Genomes, MAGs)的系统发育树,所纳入的序列还包括Burgess等人(2012)此前鉴定、完整基因组可于美国国家生物技术信息中心(NCBI,2017年9月6日下载)获取的所有分类群,以及Hug等人(2016)研究中使用的全部古菌与细菌基因组。 本系统发育树所用基因组与映射文件均可于figshare平台获取。 (Hug等人2016年生命之树所用基因组:https://doi.org/10.6084/m9.figshare.6863594.v1、https://doi.org/10.6084/m9.figshare.6863744.v2、https://doi.org/10.6084/m9.figshare.6863813.v1;Burgess等人(2012)所用基因组:https://doi.org/10.6084/m9.figshare.6863798.v1) PhyloSift工具可针对每个分类群的一套核心标记基因生成拼接序列比对。本次研究选用其中37个单拷贝标记基因,包括:核糖体蛋白S2(rpsB)、S10(rpsJ)、L1(rplA)、L22、L4/L1e(rplD)、L2(rplB)、S9(rpsl)、L3(rplC)、L14b/L23e(rplN)、S5、S19(rpsS)、S7、L16/L10E(rplP)、S13(rpsM)、L15、L25/L23、L6(rplF)、L11(rplK)、L5(rplE)、S12/S23、L29、S3(rpsC)、S11(rpsK)、L10、S8、L18P/L5E、S15P/S13e、S17、S13(rplM)、L24;以及翻译起始因子IF-2、金属内肽酶、苯丙氨酰-tRNA合成酶β亚基、苯丙氨酰-tRNA合成酶α亚基、tRNA假尿苷合酶B、胆色素原脱氨酶与核糖核酸酶HII;即选取PhyloSift标记DNGNGWU00001至DNGNGWU00040中排除DNGNGWU00004、DNGNGWU00008与DNGNGWU00038的序列。 针对上述37个拼接基因的氨基酸序列比对,我们使用trimAl v.1.2进行了修剪:移除了在超过5%的序列中存在缺失的比对列,同时剔除了覆盖拼接序列比例低于75%的分类群。不过来自ARK与ZAV的宏基因组组装基因组未满足该阈值,我们将其手动保留于本次比对中。

创建时间:
2018-08-09
二维码
社区交流群
二维码
科研交流群
商业服务