Comprehensive survey of conserved RNA secondary structures in full-genome alignment of Hepatitis C virus - Supplementary Files
收藏资源简介:
This is the repository for all supplementary files of the publication: Triebel, S., Lamkiewicz, K., Ontiveros, N. et al. Comprehensive survey of conserved RNA secondary structures in full-genome alignment of Hepatitis C virus. Sci Rep 14, 15145 (2024). https://doi.org/10.1038/s41598-024-62897-0 Abstract Hepatitis C virus (HCV) is a plus-stranded RNA virus that often chronically infects liver hepatocytes and causes liver cirrhosis and cancer. These viruses replicate their genomes employing error-prone replicases. Thereby, they routinely generate a large ‘cloud’ of RNA genomes (quasispecies) which - by trial and error - comprehensively explore the sequence space available for functional RNA genomes that maintain the ability for efficient replication and immune escape. In this context, it is important to identify which RNA secondary structures in the sequence space of the HCV genome are conserved, likely due to functional requirements. Here, we provide the first genome-wide multiple sequence alignment (MSA) with the prediction of RNA secondary structures throughout all representative full-length HCV genomes. We selected 57 representative genomes by clustering all complete HCV genomes from the BV-BRC database based on k-mer distributions and dimension reduction and adding RefSeq sequences. We include annotations of previously recognized features for easy comparison to other studies. Our results indicate that mainly the core coding region, the C-terminal NS5A region, and the NS5B region contain secondary structure elements that are conserved beyond coding sequence requirements, indicating functionality on the RNA level. In contrast, the genome regions in between contain less highly conserved structures. The results provide a complete description of all conserved RNA secondary structures and make clear that functionally important RNA secondary structures are present in certain HCV genome regions but are largely absent from other regions. Full-genome alignments of all branches of Hepacivirus C are provided in the supplement. The supplementary files include: F1 - original data set in Fasta (fasta) format F2 - pre-filtered data set in Fasta (fasta) format F3 - phylogenetic tree in Newick (treefile) and Nexus (nex) format F4 - representative genomes in Fasta (fasta) format F5 - nucleotide alignment in ClustalW (clustal), Fasta (fasta), and Stockholm (stk) format (latterformat with additional annotations such as RNA secondary structures and genes) F6 - protein alignment in ClustalW (clustal), Fasta (fasta), and Stockholm (stk) format F7 - nucleotide and protein alignment in Stockholm (stk) format (with additional annotations such as RNAsecondary structures and genes) F8 - alignment of the 5’ UTR (51 sequences) in Stockholm (stk), ClustalW (clustal), and Fasta (fasta)format F9 - alignment of the 3’ X-tail (11 sequences) in Stockholm (stk), ClustalW (clustal), and Fasta(fasta) format
本仓库收录了下述学术论文的全部补充文件: Triebel, S., Lamkiewicz, K., Ontiveros, N. 等. 丙型肝炎病毒全基因组比对中保守RNA二级结构的全面综述. Sci Rep 14, 15145 (2024). https://doi.org/10.1038/s41598-024-62897-0 摘要 丙型肝炎病毒(Hepatitis C virus, HCV)是一种正链RNA病毒,通常会慢性感染肝脏肝细胞,并引发肝硬化与肝癌。这类病毒使用保真性较低的复制酶进行基因组复制,由此会持续产生大量RNA基因组集群(准物种,quasispecies)——通过试错机制,它们能够全面探索可形成功能性RNA基因组的序列空间,这类基因组需维持高效复制与免疫逃逸的能力。在此背景下,鉴定HCV基因组序列空间中哪些RNA二级结构因功能需求而保守,具有重要意义。 本研究首次提供了全基因组多序列比对(Multiple Sequence Alignment, MSA),并对所有代表性全长HCV基因组的RNA二级结构进行了预测。我们通过基于k-mer分布与降维聚类的方法,从BV-BRC数据库中提取所有完整HCV基因组,并加入RefSeq序列,最终筛选出57个代表性基因组。为便于与其他研究对比,本研究包含了已被证实的基因组特征注释。 研究结果表明,核心编码区、C端NS5A区以及NS5B区主要包含超出编码序列需求的保守二级结构元件,提示这些结构在RNA层面发挥功能。与之相反,区间内的基因组区域则含有较少高度保守的结构。本研究结果完整描述了所有保守RNA二级结构,并明确了具有功能重要性的RNA二级结构仅存在于HCV基因组的特定区域,而在其他区域基本缺失。补充材料中提供了丙型肝炎病毒(Hepacivirus C)所有分支的全基因组比对结果。 补充文件包含: F1 - Fasta(fasta)格式的原始数据集 F2 - Fasta(fasta)格式的预过滤数据集 F3 - Newick(treefile)与Nexus(nex)格式的系统发育树 F4 - Fasta(fasta)格式的代表性基因组序列 F5 - ClustalW(clustal)、Fasta(fasta)与Stockholm(stk)格式的核苷酸比对文件(后一种格式包含RNA二级结构与基因等额外注释) F6 - ClustalW(clustal)、Fasta(fasta)与Stockholm(stk)格式的蛋白质比对文件 F7 - Stockholm(stk)格式的核苷酸与蛋白质比对文件(包含RNA二级结构与基因等额外注释) F8 - 5’非翻译区(5’ UTR,共51条序列)的比对文件,格式为Stockholm(stk)、ClustalW(clustal)与Fasta(fasta) F9 - 3’ X-tail区(共11条序列)的比对文件,格式为Stockholm(stk)、ClustalW(clustal)与Fasta(fasta)



