A transformer model for de novo sequencing of data-independent acquisition mass spectrometry data
收藏资源简介:
A core computational challenge in the analysis of mass spectrometry data is the de novo sequencing problem, in which the generating amino acid sequence is inferred directly from an observed fragmentation spectrum without the use of a sequence database. Recently, deep learning models have made significant advances in de novo sequencing by learning from massive datasets of high confidence labeled mass spectra. However, these methods are primarily designed for data-dependent acquisition (DDA) experiments. Over the past decade, the field of mass spectrometry has been moving toward using data-independent acquisition (DIA) protocols for the analysis of complex proteomic samples due to their superior specificity and reproducibility. Hence, we present a new de novo sequencing model called Cascadia, which uses a transformer architecture to handle the more complex data generated by DIA protocols. In comparisons with existing approaches for de novo sequencing of DIA data, Cascadia achieves improved performance across a range of instruments and experimental protocols. Additionally, we demonstrate Cascadia’s ability to accurately discover de novo coding variants and peptides from the variable region of antibodies.
质谱数据分析中的核心计算挑战为从头测序(de novo sequencing)问题,即无需借助序列数据库,直接从观测到的肽段碎裂谱图中推断出目标氨基酸序列。近年来,深度学习模型通过在海量高置信度标注质谱数据集上进行学习,在从头测序任务中取得了显著进展。然而,此类方法主要针对数据依赖采集(data-dependent acquisition, DDA)实验设计。近十年来,由于数据非依赖采集(data-independent acquisition, DIA)协议具备更优异的特异性与重现性,质谱领域正逐步转向使用该协议分析复杂蛋白质组样本。 为此,我们提出一款全新的从头测序模型Cascadia,其采用Transformer架构以处理DIA协议产生的更为复杂的实验数据。在与现有DIA数据从头测序方法的对比实验中,Cascadia在多种仪器与实验协议下均实现了更优的性能表现。此外,我们还验证了Cascadia可精准发现抗体可变区中的从头编码变异体与肽段。



