High resolution annotation of Zebrafish transcriptome using long-read sequencing
收藏资源简介:
With the emergence of zebrafish as an important model organism, a concerted effort has been made to study its transcriptome. This effort is limited by gaps in zebrafish annotation, which is especially pronounced concerning transcripts dynamically expressed during zygotic genome activation (ZGA). To date, short read sequencing has been the principal technology for zebrafish transcriptome annotation. In part because these sequence reads are too short for assembly methods to resolve the full complexity of the transcriptome, the current annotation is rudimentary. By providing direct observation of full-length transcripts, recently refined long-read sequencing platforms can dramatically improve annotation coverage and accuracy. Here, we leveraged the SMRT platform to study the early ZGA-stage zebrafish transcriptome. Our analysis revealed additional novelty and complexity in the zebrafish transcriptome, identifying 2748 high confidence novel transcripts that originated from previously unannotated loci and 1835 new isoforms in previously annotated genes. Overall design: Pooled RNA of a-amanitin / untreated embryos were collected and profiled with long-read sequencing. Temporally corresponding pre/post ZGA pooled embryonic RNA samples were profiled with short-read RNA-seq. Long-read raw data were assembled into transcripts using IsoSeq  (PMID: 27407110), mapped to the reference GRCz10 genome using GMAP [PMID:15728110] and annotated against the reference transcriptome using Cuffcompare [PMC3334321]. Novel transcripts were compared to short-read data and computationally validated in constructing a final long-read augmented transcriptome.
随着斑马鱼作为重要模式生物(model organism)的兴起,学界已开展协同攻关以解析其转录组。但该类研究受限于斑马鱼基因组注释存在的诸多缺口,尤其在合子基因组激活(zygotic genome activation, ZGA)过程中动态表达的转录本相关注释方面尤为突出。迄今为止,短读长测序(short read sequencing)一直是斑马鱼转录组注释的主流技术。部分原因在于这类序列读长过短,现有组装方法难以解析转录组的全部复杂特征,因此当前的注释结果仍较为粗糙。新近优化的长读长测序平台可直接获取全长转录本,从而大幅提升注释的覆盖度与准确性。本研究借助SMRT平台对合子基因组激活早期的斑马鱼转录组展开分析。本次分析揭示了斑马鱼转录组中更多此前未被发现的特征与复杂结构,共鉴定得到2748条高可信度的全新转录本(源自未注释的基因座)以及1835个已注释基因的新转录本亚型。实验设计:收集α-鹅膏蕈碱(a-amanitin)处理与未处理胚胎的混合RNA,通过长读长测序进行转录组分析;同时对对应时间点的合子基因组激活前后的混合胚胎RNA样本开展短读长RNA测序(short-read RNA-seq)。长读长原始数据通过IsoSeq(PMID: 27407110)组装为转录本,再通过GMAP[PMID:15728110]比对至参考基因组GRCz10,并借助Cuffcompare[PMC3334321]与参考转录组进行注释比对。将全新转录本与短读长测序数据进行比对,并通过计算验证,最终构建得到长读长测序辅助优化的转录组。



