Equus caballus Genome sequencing and assembly. Equus caballus
收藏资源简介:
The current reference assembly for the domestic horse Equus caballus was published in 2009. This assembly used Sanger (first-generation) sequencing along with bacterial artificial chromosomes (BACs) to produce an assembly with 112kb contig N50 and 46Mb scaffold N50. Since 2009, there have been enormous advances in sequencing technology, bringing the cost and time required for sequencing down by more than a factor of 10,000. We used the existing Sanger sequence data along with several of these new technologies to assemble a new reference genome. Illumina HiSeq short reads increased average assembly read depth six-fold, resulting in fewer mis-calls. We used CHiCago and Hi-C long-insert libraries to improve scaffold assembly, nearly doubling the scaffold N50 and increasing the amount of sequence assigned to chromosomes. Gap-filling with PacBio long reads greatly increased contig sizes. We used a 10x Chromium library to identify and phase variants. The final assembly has 4,493kb contig N50, 85Mb scaffold N50, and 70Mb more sequence assigned to chromosomes.
家马(Equus caballus)的现行参考基因组组装版本于2009年发表。该组装版本采用桑格(第一代)测序技术结合细菌人工染色体(BACs),最终获得的组装结果重叠群N50(contig N50)为112kb,支架N50(scaffold N50)为46Mb。自2009年以来,测序技术实现了跨越式进步,测序所需的成本与耗时均缩减至原有水平的万分之一以下。本研究结合原有桑格测序数据与多项新兴测序技术,完成了全新的家马参考基因组组装:通过引入Illumina HiSeq短读长序列,将组装的平均测序深度提升六倍,有效降低了碱基误判率;使用CHiCago与Hi-C长插入片段文库优化支架组装流程,使支架N50提升近一倍,同时增加了可锚定至染色体的序列总量;通过PacBio长读长序列进行缺口填补,大幅提升了重叠群的序列长度;此外还利用10x Chromium文库完成了变异位点的识别与分型工作。最终组装版本的重叠群N50达到4493kb,支架N50达到85Mb,可锚定至染色体的序列总量较此前版本增加了70Mb。



