Meraculous: <em>De Novo</em> Genome Assembly with Short Paired-End Reads
收藏资源简介:
We describe a new algorithm, meraculous, for whole genome assembly of deep paired-end short reads, and apply it to the assembly of a dataset of paired 75-bp Illumina reads derived from the 15.4 megabase genome of the haploid yeast Pichia stipitis. More than 95% of the genome is recovered, with no errors; half the assembled sequence is in contigs longer than 101 kilobases and in scaffolds longer than 269 kilobases. Incorporating fosmid ends recovers entire chromosomes. Meraculous relies on an efficient and conservative traversal of the subgraph of the k-mer (deBruijn) graph of oligonucleotides with unique high quality extensions in the dataset, avoiding an explicit error correction step as used in other short-read assemblers. A novel memory-efficient hashing scheme is introduced. The resulting contigs are ordered and oriented using paired reads separated by ∼280 bp or ∼3.2 kbp, and many gaps between contigs can be closed using paired-end placements. Practical issues with the dataset are described, and prospects for assembling larger genomes are discussed.
我们提出一种用于深度双端短读长序列全基因组组装的新型算法Meraculous,并将其应用于源自单倍体酵母斯氏毕赤酵母(Pichia stipitis)15.4 Mb基因组的成对75 bp Illumina测序读段数据集的组装。该算法可回收超过95%的基因组序列且无错误;组装序列的一半位于长度超过101 kb的重叠群(contig)以及长度超过269 kb的支架序列(scaffold)中。引入粘粒末端(fosmid end)信息可完整回收整条染色体。Meraculous依托对数据集中带有独特高质量延伸的寡核苷酸的k元组(k-mer)德布鲁因图(de Bruijn图)子图进行高效且保守的遍历,无需像其他短读长组装工具那样执行显式的错误校正步骤。本文提出了一种新型内存高效的哈希方案。利用间距约为280 bp或3.2 kbp的双端读段对所得重叠群进行排序与定向,并可通过双端读段的比对位置填补众多重叠群之间的缺口。本文还阐述了该数据集面临的实际问题,并探讨了组装更大基因组的应用前景。



