遇见数据集

Linked read technology for assembling large complex and polyploid genomes

收藏
NIAID Data Ecosystem2026-05-26 收录
官方服务:

资源简介:

Short read DNA sequencing technologies have revolutionized sequencing by providing high accuracy and throughput at low cost, but applications are limited due to the difficulties of assembling and identifying unique genomic locations of short reads. The linked read strategy overcomes these limitations because all short reads originating from a single long molecule of DNA share a common barcode. However, the majority of studies to date that have employed linked reads were focused on human haplotype phasing and genome assembly. Here we describe a de novo maize B73 genome assembly generated via linked read technology which contains ~172,000 scaffolds that cover 50% of the genome. Based on comparisons to the B73 reference genome, 91% of linked read contigs are accurately assembled, and errors were identified with >76% accuracy using a machine learning approach, suggesting that it may be possible to identify and potentially correct systematic errors. The linked read assembly contains substantially longer contigs than the assembly constructed without reference to the long molecule information (N50 of 14.5 kb and 238 bp, respectively) with similar proportions of the genome covered and similar accuracies. Complex polyploids represent one of the last grand challenges in genome assembly. Our results demonstrate that linked read technology can successfully resolve the two subgenomes of a recent alloployloid, i.e., proso millet (Panicum miliaceum). Our proso millet assembly covers ~83% of the 1 Gb genome and consists of 30,819 scaffolds with an N50 of 912 kb. Our analysis provides a framework for future de novo genome assemblies using linked reads, and we suggest computational strategies that if implemented have the potential to improve linked read assemblies, particularly for repetitive genomes.

短读长DNA测序技术(short read DNA sequencing technologies)凭借低成本实现高精度与高通量测序,彻底变革了测序领域,但由于难以组装和识别短读长的独特基因组位置,其应用受到了限制。连接读长技术(linked read)可克服上述局限,因为源自单条DNA长分子的所有短读长共享同一条形码(barcode)。然而迄今为止,多数采用连接读长技术的研究均聚焦于人类单倍型分型(haplotype phasing)与基因组组装。本研究介绍了一种通过连接读长技术构建的玉米B73从头(de novo)基因组组装结果,该组装包含约172000个支架序列(scaffold),覆盖50%的基因组。通过与B73参考基因组比对,91%的连接读长重叠群(contig)组装准确;采用机器学习方法可识别出其中76%以上的错误,这表明识别并潜在校正系统误差具备可行性。连接读长组装得到的重叠群长度显著优于未参考长分子信息构建的组装结果(二者N50值分别为14.5千碱基对(kb)与238碱基对(bp)),且二者覆盖基因组的比例与组装准确度相近。复杂多倍体是基因组组装领域尚存的重大挑战之一。本研究结果表明,连接读长技术可成功解析新近形成的异源多倍体(alloployloid)的两个亚基因组,即黍(Panicum miliaceum)。我们的黍基因组组装结果覆盖1吉碱基对(Gb)基因组的约83%,包含30819个支架序列,其N50值为912 kb。本研究的分析为未来采用连接读长技术的从头基因组组装提供了框架,并提出了若干计算策略,若得以实施,有望优化连接读长组装结果,尤其针对重复基因组。

创建时间:
2018-08-17
二维码
社区交流群
二维码
科研交流群
商业服务