遇见数据集

Enhanced detection of RNA modifications with high-accuracy nanopore RNA basecalling models

收藏
官方服务:

资源简介:

Chemical RNA modifications, collectively referred to as the ‘epitranscriptome’, have been intensively studied during the last years, largely facilitated by the use of next-generation sequencing technologies. Recent efforts have turned towards the nanopore direct RNA sequencing (DRS) platform, as it allows simultaneous detection of diverse RNA modification types in full-length native RNA molecules. While RNA modifications can be identified in the form of systematic basecalling ‘errors’ in DRS datasets, m6A modifications produce very modest ‘errors’, limiting the applicability of this approach to sites that are modified at high stoichiometries. Here, we demonstrate that the use of alternative RNA basecalling models, trained with fully-unmodified in vitro synthetic sequences, increase the ‘error’ signal of m6A modifications, leading to enhanced detection of RNA modifications even at lower stoichiometries. We then show that the use of these models enhances the detection of RNA modifications on previously published in vivo human samples, using third-party softwares for the detection of RNA modifications. Moreover, our work provides a novel RNA basecalling model that shows a median accuracy of 97%, compared to previously available RNA basecalling models that show 91% accuracy. Notably, this increase in accuracy does not only lead to improved detection of RNA modifications, but also enhanced mappability of RNA reads, which becomes more evident in the case of short RNA reads (50% increase). Altogether, our work stresses the importance of using fully unmodified RNA sequences for training RNA basecalling models, and how the use of different basecalling models can significantly affect the detection of RNA modifications and read mappability.

被统称为表观转录组(epitranscriptome)的化学性RNA修饰,在近年来得到了深入研究,这在很大程度上得益于二代测序技术(next-generation sequencing technologies)的广泛应用。近期的研究重心转向了纳米孔直接RNA测序(nanopore direct RNA sequencing, DRS)平台,因其可在全长天然RNA分子中同时检测多种RNA修饰类型。尽管在DRS数据集中,RNA修饰可通过系统性碱基识别(basecalling)‘误差’的形式被识别,但N6-甲基腺嘌呤(m6A)修饰仅能产生极微弱的‘误差’信号,这使得该方法仅适用于高修饰比例的修饰位点。本研究证实,使用以完全未修饰的体外(in vitro)合成RNA序列训练得到的替代型RNA碱基识别模型,可增强m6A修饰的‘误差’信号,即便在较低修饰比例的位点也能提升RNA修饰的检测效能。随后,本研究借助第三方RNA修饰检测软件,证明了此类模型可优化已发表的体内(in vivo)人类样本的RNA修饰检测效果。此外,本研究构建了一款新型RNA碱基识别模型,其中位数准确率可达97%,而此前已有的RNA碱基识别模型的准确率仅为91%。值得注意的是,准确率的提升不仅优化了RNA修饰的检测效果,还增强了RNA读段的比对效能,这一优势在短RNA读段中尤为显著(比对率提升50%)。综上,本研究强调了使用完全未修饰的RNA序列训练RNA碱基识别模型的重要性,同时也揭示了选用不同碱基识别模型可对RNA修饰检测及读段比对效能产生显著影响。

二维码
社区交流群
二维码
科研交流群
商业服务