Identifying and Classifying Trait Linked Polymorphisms in Non-Reference Species by Walking Coloured de Bruijn Graphs
收藏资源简介:
Single Nucleotide Polymorphisms are invaluable markers for tracing the genetic basis of inheritable traits and the ability to create marker libraries quickly is vital for timely identification of target genes. Next-generation sequencing makes it possible to sample a genome rapidly, but polymorphism detection relies on having a reference genome to which reads can be aligned and variants detected. We present Bubbleparse, a method for detecting variants directly from next-generation reads without a reference sequence. Bubbleparse uses the de Bruijn graph implementation in the Cortex framework as a basis and allows the user to identify bubbles in these graphs that represent polymorphisms, quickly, easily and sensitively. We show that the Bubbleparse algorithm is sensitive and can detect many polymorphisms quickly and that it performs well when compared with polymorphism detection methods based on alignment to a reference in Arabidopsis thaliana. We show that the heuristic can be used to maximise the number of true polymorphisms returned, and with a proof-of-principle experiment show that Bubbleparse is very effective on data from unsequenced wild relatives of potato and enabled us to identify disease resistance linked genes quickly and easily.
单核苷酸多态性(Single Nucleotide Polymorphisms)是追踪遗传性状遗传基础的极为宝贵的分子标记,而快速构建标记库对于及时鉴定靶基因至关重要。下一代测序(Next-generation sequencing)技术使得快速获取基因组样本成为可能,但多态性检测通常需要依赖参考基因组,才能将测序读段(reads)比对至该参考序列并检测变异。我们提出Bubbleparse——一种无需参考序列即可直接从下一代测序读段中检测变异的方法。该方法以Cortex框架中的德布鲁因图(de Bruijn graph)实现为基础,能够让用户快速、便捷且灵敏地识别这些代表多态性的德布鲁因图气泡。我们证实,Bubbleparse算法灵敏度优异,可快速检测出大量多态性;在拟南芥(Arabidopsis thaliana)样本中,与基于参考基因组比对的多态性检测方法相比,其表现更为出色。此外,我们证明该启发式策略可用于最大化返回的真实多态性数量;通过原理验证实验,我们证实Bubbleparse在未完成测序的马铃薯野生近缘种的测序数据中效果极佳,能够帮助研究人员快速便捷地鉴定与抗病性相关的基因。




