Supplementary Data for Kogay et al. (2019)
收藏资源简介:
This dataset contains supplementary figures and tables, sequence alignments, phylogenetic trees and outlier removal calculations used in the bioinformatic analyses presented in: Roman Kogay, Taylor B. Neely, Daniel P. Birnbaum, Camille R. Hankel, Migun Shakya, and Olga Zhaxybayeva. “Machine-learning classification suggests that many alphaproteobacterial prophages may instead be gene transfer agents”, BioRxiv, 2019. (BIORXIV/2019/697243; available at https://www.biorxiv.org/content/10.1101/697243v1) File Contents: Supplementary_Figures.pdf: Supplementary Figures S1 and S2 in the manuscript. Supplementary_Tables.zip: Supplementary Tables S1-S15 in the manuscript. Alignments_for_weight_assignment_GTAs.zip: Alignments of 'true GTA' sequences in the training dataset. These alignments were used to generate pairwise phylogenetic distances for the weighting scheme. The alignments are in FASTA format. The filename prefix (g2, …, g15) refers to the RcGTA gene name (see Supplementary Table S1). Alignments_for_weight_assignment_viruses.zip: Alignments of 'true virus' sequences in the training dataset. These alignments were used to generate pairwise phylogenetic distances for the weighting scheme. The alignments are in FASTA format. The filename prefix (g2, …, g15) refers to the RcGTA gene name (see Supplementary Table S1). Alignments_for_removal_of_outliers.zip: Alignments of ‘true GTA’ and ‘true virus’ sequences in the training datasets. Pairwise phylogenetic distances calculated from these alignments were used to remove GTA homologs that are more closely related to viruses than to other GTAs, as well as to investigate obtained lower accuracies for g6 and g12 (Supplementary Table S11). The alignments are in FASTA format. The filename prefix (g2, …, g15) refers to the RcGTA gene name (see Supplementary Table S1). Outlier_removal.xlsx: Calculations to identify GTAs that are more closely related to viruses than to other GTAs. The removed sequences are highlighted. Reference_phylogenetic_tree_reconstruction.zip: Concatenated alignment of 83 marker genes in 1,423 taxa in PHYLIP format (concatenated_83markers.phy); information about partitions and substitution models used in phylogenetic reconstruction (partitions.txt); and phylogenetic tree in Newick format (1423_alphaproteobacteria_reference_tree.newick).
本数据集包含对应文献的生物信息学分析中所用的补充图表、序列比对文件、系统发育树及异常值去除计算材料,对应文献为:Roman Kogay、Taylor B. Neely、Daniel P. Birnbaum、Camille R. Hankel、Migun Shakya与Olga Zhaxybayeva于2019年发布在BioRxiv的论文"Machine-learning classification suggests that many alphaproteobacterial prophages may instead be gene transfer agents"(预印本编号BIORXIV/2019/697243,可通过https://www.biorxiv.org/content/10.1101/697243v1获取)。 文件内容如下: 1. Supplementary_Figures.pdf:论文中的补充图S1与S2。 2. Supplementary_Tables.zip:论文中的补充表S1至S15。 3. Alignments_for_weight_assignment_GTAs.zip:训练数据集内"真正基因转移剂(gene transfer agent, GTA)"序列的比对文件,用于生成加权方案所需的成对系统发育距离,比对格式为FASTA。文件名前缀(g2、……、g15)对应RcGTA基因名称(详见补充表S1)。 4. Alignments_for_weight_assignment_viruses.zip:训练数据集内"真正病毒"序列的比对文件,用于生成加权方案所需的成对系统发育距离,比对格式为FASTA。文件名前缀(g2、……、g15)对应RcGTA基因名称(详见补充表S1)。 5. Alignments_for_removal_of_outliers.zip:训练数据集内"真正GTA"与"真正病毒"序列的比对文件,基于这些比对得到的成对系统发育距离,可用于去除与病毒亲缘关系较近而非与其他GTA亲缘关系更近的GTA同源序列,同时可用于分析g6与g12基因分类精度较低的原因(详见补充表S11)。比对格式为FASTA,文件名前缀(g2、……、g15)对应RcGTA基因名称(详见补充表S1)。 6. Outlier_removal.xlsx:用于鉴定与病毒亲缘关系较近而非与其他GTA亲缘关系更近的GTA的计算文档,已移除的序列已做高亮标注。 7. Reference_phylogenetic_tree_reconstruction.zip:包含1423个分类单元的83个标记基因的串联比对文件(格式为PHYLIP,文件名为concatenated_83markers.phy)、系统发育重建所用的分区信息与替换模型说明文件(partitions.txt),以及Newick格式的系统发育树文件(1423_alphaproteobacteria_reference_tree.newick)。




