Additional file 1: Table S1. of Spherical: an iterative workflow for assembling metagenomic datasets
收藏资源简介:
Effect of kmer size on assembly. For each dataset (column 1) the alignment rate (%) (column 3) of raw data aligning back to assembly produced using different kmers (column 2) was assessed. Table S2. Effect of kmer size across iterations of assembly using the simulated dataset. The percentage of raw data aligning to each iterations assembly was identified and iterations stopped once the alignment rate was under 0.1%. Table S3. Analysis of the quality of contigs produced in each iteration of assembling the simulated dataset using contig scores. The contig score identifies the percentage accuracy of a contig compared to the genomes used to create the simulated metagenome. We present the percentage of contigs from each iterations assembly with contig scores gereater than 95 or less than 50. Table S4. Percentage of reads assigned to each of the 400 genomes within the simulated dataset, base assembly and Spherical assembly of the simulated dataset. Table S5. Assembly statistics comparing dataset assemblies for each method. The first column indicates the dataset utilized whilst the second column identified the assembly methodology. Due to Spherical having a sub-sampling option the size of the sample utilized by Spherical was stated for each assembly in column 5. The final 6 columns provide information on the computational needs for each assembly (RAM usage) as well as statistics about the produced assemblies e.g. number of contigs and alignment (%). Table S6. The number of reads aligning to genes within the Spherical iterations assembling each dataset. Figure S1. The taxonomic variations at the phylum level between each experimental assembly method for each dataset. Each bar represents the number of reads that could be assigned to a taxonomic Phyla within each assembly method for the datasets. The legend identifies which Phyla is represented by each colour. (ZIP 562 kb)
k值(k-mer)大小对基因组组装的影响。针对每个数据集(第1列),本研究评估了原始测序数据与使用不同k值(第2列)组装得到的序列的比对率(%,第3列)。 表S2 模拟数据集组装迭代过程中k值大小的影响。本研究统计了原始测序数据与每一轮组装结果的比对率,并在比对率低于0.1%时停止迭代。 表S3 基于重叠群得分的模拟数据集迭代组装结果的重叠群质量分析。重叠群得分用于衡量单条重叠群与用于构建模拟宏基因组的参考基因组之间的序列匹配准确率百分比。本研究统计了每一轮组装结果中,重叠群得分高于95或低于50的重叠群所占百分比。 表S4 模拟数据集、该数据集的基础组装结果以及Spherical组装结果中,分配至模拟数据集内400个基因组的测序读段所占百分比。 表S5 不同方法组装数据集的统计结果对比。第1列为所用数据集,第2列为组装方法。由于Spherical支持子采样选项,因此第5列标注了Spherical每次组装时采用的样本量。最后6列则提供了每次组装的计算资源需求(如随机存取内存(RAM)使用量)以及组装结果的相关统计指标,例如重叠群数量与比对率(%)。 表S6 各数据集经Spherical迭代组装时,比对至基因区域的测序读段数量。 图S1 各数据集采用不同实验组装方法后,物种分类学层面的门水平差异。每一个条形代表对应组装方法的数据集可被分配至某一分类门的读段数量,图例标注了每种颜色所代表的分类门。(压缩包大小:562 KB)



