Sars-escape network for escape prediction of SARS-COV-2
收藏资源简介:
This dataset belongs to a paper Sars-escape network for escape prediction of SARS-COV-2 Prem Singh Bist, Hilal Tayara, Kil To Chong <em>Briefings in Bioinformatics</em>, Volume 24, Issue 3, May 2023, bbad140, https://doi.org/10.1093/bib/bbad140 Abstract <strong>Motivation: </strong>Viruses have coevolved with their hosts for over millions of years and learned to escape the host's immune system. Although not all genetic changes in viruses are deleterious, some significant mutations lead to the escape of neutralizing antibodies and weaken the immune system, which increases infectivity and transmissibility, thereby impeding the development of antiviral drugs or vaccines. Accurate and reliable identification of viral escape mutational sequences could be a good indicator for therapeutic design. We developed a computational model that recognizes significant mutational sequences based on escape feature identification using natural language processing along with prior knowledge of experimentally validated escape mutants. <strong>Results: </strong>Our machine learning-based computational approach can recognize the significant spike protein sequences of severe acute respiratory syndrome coronavirus 2 using sequence data alone. This modelling approach can be applied to other viruses, such as influenza, monkeypox and HIV using knowledge of escape mutants and relevant protein sequence datasets. <strong>Availability: </strong>Complete source code and pre-trained models for escape prediction of severe acute respiratory syndrome coronavirus 2 protein sequences are available on Github at https://github.com/PremSinghBist/Sars-CoV-2-Escape-Model.git. <strong>Contact: </strong>premsing212@jbnu.ac.kr. <strong>Keywords: </strong>SARS-CoV-2; mutation; sequence analysis; viral escape prediction.
本数据集关联于一篇题为《用于新型冠状病毒(SARS-CoV-2)逃逸预测的Sars-escape网络》的学术论文,作者为Prem Singh Bist、Hilal Tayara、Kil To Chong,发表于<em>Briefings in Bioinformatics</em> 2023年5月第24卷第3期,文章编号bbad140,DOI链接:https://doi.org/10.1093/bib/bbad140。<br><br><strong>研究背景(Motivation):</strong>病毒与宿主已共同演化数百万年,演化出逃避宿主免疫系统的能力。尽管多数病毒遗传变异无有害影响,但部分关键突变可导致病毒逃逸中和抗体,削弱宿主免疫系统,进而提升病毒感染性与传播能力,阻碍抗病毒药物或疫苗的研发。精准可靠地鉴定病毒逃逸突变序列,可为治疗方案设计提供重要参考。本研究开发了一款计算模型,该模型基于自然语言处理的逃逸特征识别,并结合经实验验证的逃逸突变体先验知识,以识别关键突变序列。<br><br><strong>研究结果(Results):</strong>本研究提出的基于机器学习的计算方法仅依靠序列数据,即可准确识别新型冠状病毒刺突蛋白的关键序列。该建模方法还可结合逃逸突变体知识与相关蛋白序列数据集,应用于流感、猴痘、HIV等其他病毒的相关研究。<br><br><strong>数据可用性(Availability):</strong>本研究针对新型冠状病毒蛋白序列逃逸预测的完整源代码与预训练模型,已上传至GitHub开源平台,链接为:https://github.com/PremSinghBist/Sars-CoV-2-Escape-Model.git。<br><br><strong>联系方式(Contact):</strong>premsing212@jbnu.ac.kr。<br><br><strong>关键词(Keywords):</strong>新型冠状病毒(SARS-CoV-2);突变;序列分析;病毒逃逸预测。



