Sequence variation in HIV-1 protease and reverse transcriptase genes.
收藏资源简介:
Sequence variation at the nucleotide, dinucleotide, codon and amino acid level was assessed in an alignment of 4455 HIV-1 sequences spanning 1017 nucleotides from codon 1 of PR to codon 240 of RT, and representing three major subtypes of the HIV-1 M group (A, B and C). Constraint and polymorphism at the nucleotide, dinucleotide, codon and amino acid level amongst wild-type (i.e. not drug selected) sequences, was evaluated separately for each subtype. Variation was assessed by scoring the frequency of mutations (defined as changes relative to the subtype majority consensus sequence) at all sites across the alignment. Plotting the distribution of nucleotide and amino acid variability reveals a picture of remarkably pervasive constraint at both levels. At the amino acid level, marked conservation reflects strong purifying selection acting on the PR and RT enzymes. Conservation at the amino acid level is reflected at the nucleic acid level, as many nucleotide sites (~55%) occur at non-synonymous positions within conserved amino acid sites. However, synonymous positions are also relatively invariable. Patterns of sequence variation were further characterised by mapping the distribution of conserved and polymorphic sites across the aligned region for each subtype. Amino acid sites were categorised as either conserved (<1% variation), constrained (>1% but <10% variation), or polymorphic (>10% variation). Nucleotide sites were categorised as synonymous or non-synonymous based on the consensus amino acid sequence. Synonymous sites were then further categorised as follows: (i) constrained typical (<10% variation from consensus and representing preferred codon usage for HIV-1, (ii) constrained atypical (<10% variation and representing atypical codon usage for HIV-1, and (iii) polymorphic (>10% variation). Mapping revealed a broadly similar distribution of conserved and polymorphic amino acid sites across all subtypes. At the nucleotide level, conservation at synonymous sites in the alignment largely reflects the pronounced G-A bias. However, a proportion of conserved synonymous nucleotide positions in each subtype (12-13%) do not exhibit typical biases. These conserved, ‘atypical’ synonymous sites are distributed throughout PR and RT, with concentrations occurring within some regions, such as the N-terminus of PR. Strikingly, the distribution of constrained atypical and/or polymorphic sites relative to constrained typical sites is well conserved. In the figure, the location of the PR and RT genes within the pol coding domain of an integrated HIV-1 provirus is shown. The PR-RT coding region is described at several levels of detail. For each gene region, the uppermost level of blocks shows the consensus amino acid sequence of subtype B, with differences in subtypes A and C highlighted beneath. The correspondence of the amino acids sequence to the established domains of the encoded proteins is indicated. The middle level of coloured blocks shows the distribution of conserved and polymorphic amino acids sites in each subtype, according to the amino acid variation key. Lollipops attached to this level indicate the location of drug resistance-associated mutations. Mutations identified as showing significant associations with treatment in this analysis are shown as closed circles, other well-characterised resistance mutations are shown as open circles. The bottom level of smaller coloured blocks shows the distribution of conserved and polymorphic nucleotide sites across the amplified region according to the nucleic acid variation key, and as discussed in the text. Predicted stem loops structures are indicated above this level, as arcs connecting sites consistently predicted as being involved in local base-pairing interactions. For the purposes of clarity, only stem loops identified in two or more subtypes are shown, and these are referenced by numbers corresponding to their details as given in the associated table (available on FigShare). Amino acid coordinates relative to each protein are shown across the top, and nucleotide coordinates relative to the entire analysed region are shown across the bottom. The location of miscellaneous genomic features discussed in the text (poly-lysine and trans-frame regions, and the RT active site) is highlighted
本研究针对4455条人类免疫缺陷病毒1型(HIV-1)序列的比对结果,分析了核苷酸、二核苷酸、密码子及氨基酸水平上的序列变异;该序列比对覆盖了从蛋白酶(Protease,PR)第1位密码子至逆转录酶(Reverse Transcriptase,RT)第240位密码子的1017个核苷酸区域,涵盖HIV-1 M组的3个主要亚型(A、B、C)。研究针对野生型(即未经药物选择)序列,分别在核苷酸、二核苷酸、密码子及氨基酸水平上评估了选择约束性与多态性,并按不同亚型单独开展分析。变异分析通过对序列比对中所有位点的突变频率进行计分完成,此处突变定义为相对于该亚型主流共有序列的碱基/氨基酸改变。绘制核苷酸与氨基酸变异的分布图谱可见,两类水平上均存在广泛的序列约束现象。在氨基酸水平上,显著的序列保守性反映出PR与RT酶受到了强烈的净化选择作用。氨基酸水平的保守性在核酸水平上同样有所体现:约55%的核苷酸位点位于保守氨基酸位点内的非同义突变位点处。但同义突变位点同样相对保守。 本研究进一步通过绘制各亚型比对区域内保守位点与多态位点的分布图谱,对序列变异模式进行了表征。氨基酸位点被分为三类:保守位点(变异率<1%)、约束位点(变异率1%~10%)及多态位点(变异率>10%)。核苷酸位点则根据共有氨基酸序列分为同义突变位点与非同义突变位点。同义位点进一步被分为三类:(i) 典型约束同义位点(与共有序列的变异率<10%,且符合HIV-1的偏好密码子使用模式);(ii) 非典型约束同义位点(与共有序列的变异率<10%,但不符合HIV-1的偏好密码子使用模式);(iii) 多态同义位点(变异率>10%)。位点分布图谱显示,所有亚型的保守与多态氨基酸位点的分布模式整体相似。在核苷酸水平上,比对序列中同义位点的保守性在很大程度上反映出显著的G-A碱基偏好性。但各亚型中仍有12%~13%的保守同义核苷酸位点未表现出典型的碱基偏好性。这些保守的“非典型”同义位点广泛分布于PR与RT区域内,且在部分区域存在聚集现象,例如PR的N端。值得注意的是,非典型约束位点及/或多态位点相对于典型约束位点的分布模式高度保守。 本图展示了整合型HIV-1前病毒的pol基因编码区内PR与RT基因的位置。PR-RT编码区以多个细节层级进行展示。在每个基因区域中,最上方的色块区块展示了B亚型的共有氨基酸序列,A、C亚型与该序列的差异在其下方标注高亮。图中还标注了氨基酸序列与编码蛋白已知结构域的对应关系。中间层的彩色色块区块根据氨基酸变异注释表,展示了各亚型中保守与多态氨基酸位点的分布情况。附着于该层的棒棒糖状标记指示了与药物抗性相关的突变位点位置。本分析中鉴定出与治疗应答存在显著关联的突变为实心圆点,其余已被充分表征的抗性突变为空心圆点。最下方的小型彩色色块区块根据核酸变异注释表及正文所述,展示了扩增区域内保守与多态核苷酸位点的分布情况。该层上方以弧线标注了预测的茎环结构,弧线连接了被一致预测参与局部碱基配对相互作用的位点。为保证图表清晰,仅展示了在2个及以上亚型中被鉴定出的茎环结构,这些结构通过编号对应其在关联表格(可于FigShare获取)中的详细信息。图表顶部标注了相对于各蛋白的氨基酸坐标,底部标注了相对于整个分析区域的核苷酸坐标。正文所述的其他基因组特征(多聚赖氨酸区域、移码区域及RT活性位点)的位置也在图中进行了高亮标注。



