VP24, VP35, and Glycoprotein Mutation Burden Across 28 Public Zaire ebolavirus Genomes Span=2014 - 2023
收藏资源简介:
This dataset documents a systematic attempt to perform protein-level mutation burden analysis on three key Zaire ebolavirus proteins such as VP24, VP35, and glycoprotein (GP) using 28 publicly available genome assemblies from global surveillance efforts (2014 - 2023). Despite a robust computational pipeline designed to extract coding sequences (CDS), translate, and align to reference proteins (AHX24653.1, AAD14582.1, AAD14585.1), consistent extraction of full-length VP24 and GP was not possible due to pervasive 5′-end truncation in non-reference submissions.All non-KJ660346 genomes exhibited a uniform 84-nucleotide deletion at the 5′ terminus, shifting genomic coordinates and rendering standard CDS annotations invalid. While VP35 (located near the 5′ end) could be partially analyzed yielding artifactual ~94% divergence when misaligned VP24 and GP extraction failed across all truncated genomes, despite their biological presence in the sequence data. Only a direct comparison of two full-length protein sequences (Makona 2014 vs. Mayinga 1976) revealed a single conservative mutation (K212M in VP24), visualized structurally using PDB 5F1B.This dataset includes:Raw mutation burden TSVs (including failed extractions),Per-sequence FASTA and QC reports,Structural visualizations / images (PyMOL etc.),Metadata mappings and coordinate logs.We openly share these results not as a successful analysis, but as a cautionary audit: global genomic surveillance data, while abundant, often lacks the completeness and annotation consistency required for cross-isolate protein-level inference. Reproducible evolutionary virology demands not just more sequences, but better-structured, full-length, and accurately annotated genomes.No biological interpretation is provided. All outputs reflect direct computational results using best-practice tools under documented data constraints.Study by: TahirHB@GVAtlas.Org
本数据集记录了一项系统性研究,针对扎伊尔埃博拉病毒(Zaire ebolavirus)的VP24、VP35及糖蛋白(glycoprotein, GP)这三种关键蛋白开展蛋白层面的突变负荷分析,所用数据来自2014年至2023年全球监测项目公开的28组基因组组装结果。尽管搭建了成熟的计算流程,用于提取编码序列(coding sequences, CDS)、完成翻译并与参考蛋白(AHX24653.1、AAD14582.1、AAD14585.1)进行序列比对,但由于非参考提交序列普遍存在5'端截短现象,无法稳定获取全长VP24与GP序列。所有非KJ660346的基因组均在5'端存在统一的84核苷酸缺失,这改变了基因组坐标,导致标准编码序列注释失效。虽位于5'端附近的VP35可进行部分分析,但当截短基因组无法成功提取VP24和GP时,比对错位会使VP35产生约94%的人为分化假象——尽管这些蛋白实际存在于序列数据中。仅对两条全长蛋白序列(2014年Makona株与1976年Mayinga株)的直接比对,仅发现一处保守突变(VP24的K212M),并通过PDB(Protein Data Bank, PDB)5F1B完成结构可视化。 本数据集包含以下内容: 1. 原始突变负荷制表符分隔值(Tab-Separated Values, TSV)文件(含提取失败的结果) 2. 单序列FASTA格式文件与质量控制报告 3. 结构可视化图像(基于PyMOL等工具生成) 4. 元数据映射表与坐标日志 我们公开分享这些研究结果,并非因其为成功的分析案例,而是作为一项警示性审计:全球基因组监测数据虽体量庞大,但往往缺乏跨分离株蛋白层面推断所需的完整性与注释一致性。可重复的进化病毒学研究不仅需要更多序列,更需要结构完善、全长完整且注释准确的基因组。本研究未提供任何生物学解读,所有输出均为在已记录的数据约束下,使用行业最佳实践工具得到的直接计算结果。本研究由TahirHB@GVAtlas.Org完成。




