遇见数据集

File S1 - Micro-Plasticity of Genomes As Illustrated by the Evolution of Glutathione Transferases in 12 <i>Drosophila</i> Species

收藏
NIAID Data Ecosystem2026-03-09 收录
官方服务:

资源简介:

Supplementary Figures and Tables. Figure S1. The annotated gene region of Dsim\GD17126 (DsimGSTT4) from Flybase. A. 6 exons of DsimGSTT4 code for a protein of 288 amino acids. If the putative exons 2 and 3 are intron as in DmelGSTT4, DsimGSTT4 would encode a GST protein of 237 amino acids. B. The GSTT4 amino acid alignment of the proposed D. simulans and D. melanogaster shows 100% identity. Figure S2. The annotated gene region of Dere\GG21888 (DereGSTE6) from Flybase. A. The underlined nucleotides at the 5′end is the excluded sequence which makes DereGSTE6 larger than usual for Epsilon class GSTs. Exclusion of these 54 nucleotides and starting the translation from the second ATG start site as shown in red, DereGSTE6 will code for a GST protein of 222 amino acids. B. The GSTE6 amino acid sequence alignment of D. erecta and D. melanogaster shows 93% identity. Figure S3. The annotated gene region of Dana\GF17942 from Flybase. A. A possible ATG start site is situated in the intron region as shown by the underline and then read through the second exon. The coding sequence of Dana\GF17942 will become one single exon and code for a protein of 220 amino acids. B. The amino acid sequence alignment of Dana\GF17942 and DmelGSTD2 shows 64% identity. Figure S4. The annotated gene region of Dana\GF11968 (DanaGSTE11) from Flybase. A. 45 nucleotides of the coding sequences of DanaGSTE11 including stop codon are in the intron region. Including these 45 nucleotides as the last part of exon 2, DanaGSTE11 will be a 225 GST protein. B. The GSTE11 amino acid sequence alignment of D. ananassae and D. melanogaster shows 87% identity. Figure S5. The annotated gene region of Dvir\GJ22855 from Flybase. A. The suggested ATG start site is 48 nucleotides (amino acids encoded are shown below) upstream of exon 1. This now larger exon is in-frame and contiguous and encodes a protein of 216 amino acids. B. The amino acid sequence alignment of our curated Dvir\GJ22855 and DmelGSTD2 shows 67% amino acid identity. The arrow indicates the previously annotated transcriptional start site methionine (M). Figure S6. The annotated gene region of Dvir\GJ22856 from Flybase. A. The possible ATG start site is 33 nucleotides (amino acids encoded are shown below) upstream of exon 1. This larger exon is in-frame and contiguous and encodes a protein of 211 amino acids. B. The amino acids sequence alignment of Dvir\GJ22856 and DmelGSTD2 shows 64% amino acid identity. The arrow indicates the previous annotated methionine (M). Figure S7. The annotated gene region of Dvir\GJ23571 (DvirGSTZ1) from Flybase. A. The possible ATG start site is 255 nucleotides upstream of exon 1 as shown in underline. This larger exon is in-frame and contiguous and codes for a protein of 247 amino acids. B. The GSTZ1 amino acids sequence alignment of D. virillis and D. melanogaster shows 61% identity. The arrow indicates the previous annotated methionine (M). Figure S8. The annotated gene region of Dsec\GM23038 (DsecGSTT3) from Flybase. A. The incomplete gene sequencing of D. sechellia GSTT3 shown as Ns in the middle of the gene. B. The amino acid sequence alignment of GSTT3 D. sechellia and D. melanogaster is highly conserved. Figure S9. The annotated gene region of Dper\GL26999 (DperGSTT4) from Flybase. A. We have proposed that exons 2 and 3 of annotated DperGSTT4 are intron and the remainder of the coding sequence is in the incompletely sequenced region and exons 4, 5 and 6. The sequence of exons 5 and 6 is quite conserved compared to the last 2 exons from D. melanogaster. The new sequence of DperGSTT4 and the incomplete part are shown as underline. This gene needs to be re-sequenced to obtain the missing sequence. B. The amino acid alignment of GSTT4 D. persimilis and D. melanogaster shows high identity to each other. Figure S10. The annotated gene region of Dper\GL13668 (DperGSTZ2). A. ATG start codon (as underline) of GSTZ2B, Z2C and Z2A are on exon 1, 2 and 3 respectively. They share exons 4 and 5. Exons 2 and 3 are intron of GSTZ2B, and exon 3 is intron of GSTZ2C, therefore the incomplete sequencing data does not affect the sequence translation of these isoforms. Fortunately, the coding regions of GSTZ2B and Z2C were sequenced. As their sequences are quite conserved with D. melanogaster they can be manually curated. B. Based on D. melanogaster, this gene would have 6 exons. But we found a single guanosine (G) insertion (shown in red) which causes a frameshift mutation. If the extra G is absent and the 59 nucleotides after this G are intron (double underlined) as in D. melanogaster, the DperGSTZ2A, Z2B and Z2C will show 96, 100 and 100% amino acid sequence identity to all 3 spliced products of DmelGSTZ2, respectively. Obviously, with the Ns in intron 2 and this extra G in exon 5 the gene should be sequenced again. Figure S11. The annotated gene region of Dsec\GM24019 from Flybase. A. The Dsec\GM24019 gene is located next to Dsec\GM24018 gene on the same chromosome. There is a 149 nucleotide gap between the 2 genes. Although 69 nucleotides upstream of the annotated ATG start site of Dsec\GM24019 (underline) code for GST conserved sequence, this GST is still too short be an active enzyme. B. The short amino acid sequence of Dsec\GM24019 shows 52% identity to the C-terminus part of DmelGSTD6. Figure S12. The annotated gene region of Dsec\GM21877 from Flybase. A. Gene region of Dsec\GM21877. The coding sequence is 447 bp. B. The nucleotide sequence alignment of Dsec\GM21877 and Dmel\CR43687 pseudogene. Figure S13. The annotated gene region of Dyak\GE11955 (DyakGSTE1) from Flybase. A. Dyak\GE11955 was reported to have 1 intron. However the sequence of the annotated intron codes for part of the conserved DmGSTE1 protein. B. The amino acid sequence alignment of DmGSTE1 and DyakGSTE1, (1) and (2). DyakGSTE1 (1) is the protein sequence reported by Flybase. DyakGSTE1 (2) includes the protein coding sequence in the annotated intron region. Figure S14. The annotated gene region of Dana\GF17941 from Flybase. Dana\GF17941 shows high amino acid sequence identity to the C-terminus part of DmelGSTD2, 4 and 5. Figure S15. The annotated gene region of Dper\17151 (DperGSTT2) from Flybase. A. Gene region of DperGSTT2 which codes for protein 212 amino acids. The sequence show high conservative to DmelGSTT2. B. The amino acid sequence alignment of DmelGSTT2 and the pseudogene product DperGSTT2. Figure S16. The annotated gene region of Dper\GL26929 from Flybase. A. The gene region of Dper\GL26929 codes for a protein of 163 amino acids. B. The amino acid sequence alignment of Dper\GL26929 to DmelGSTT1 and T2 show 27% and 29% amino acid identity. The conserved sequences are shown underlined. Figure S17. The annotated gene region of Dwil\GK11203 from Flybase. A. Exon 1 is the annotated coding sequence from Flybase but we found conserved blocks of sequences located upstream of the annotated ATG start site as shown in the bracket. B. Dwil\GK11203 (1) is the annotated sequence reported by Flybase. It shows the highest amino acid sequence identity to DmelGSTD5 but it has about 60 critical amino acids missing from the active site so would not be an active GST enzyme. Therefore this gene is suggested to be a pseudogene. Dwil\GK11203 (2) shows some conserved sequence upstream of the annotated ATG start site. Figure S18. The annotated gene region of Dvir\GJ24387 (DvirGSTD10) from Flybase. A. Dvir\GJ24387has been reported to have 2 exons and 1 intron. Although we found that part of the intron sequence (underlined) also codes for conserved sequence of DmGSTD10, this protein is still too short to be an active GST enzyme. The conserved sequence is QYGKDSTLYPKDIQTQALIN. B. The short amino acid sequence of DvirGSTD10 shows 49% amino acid sequence identity to DmelGSTD10. Figure S19. The annotated gene region of Dvir\GJ19066 from Flybase. A. Dvir\GJ19066 encodes a protein of 54 amino acids which shows the greatest amino acid identity to DmelGSTD1. Moreover we found possible conserved sequence of 63 nucleotides upstream contiguous with the annotated ATG start site (underlined). B. The short amino acid sequence of Dvir\GJ19066 shows 20% identity to the N-terminus part of DmelGSTD1. Figure S20. The annotated gene region of Dsim\GD17492 (DsimGSTT3) from Flybase. A. The underlined ATG (exon 1 and exon 4) represent the start codons of T3B and T3A, respectively. The sequence from exons 5 to 6 of the DsimGSTT3 gene comprises the sequence of exon 5 in D. melanogaster. Due to the first cytosine (C) of exon 6 is missing; the annotation split the sequence into 2 exons apparently to keep the remaining sequence in frame. Thus DsimGSTT3A and T3B are shorter than T3A and T3B in other species. B. DsimT3 (1) refers to the sequence reported by Flybase. DsimT3 (2) refers to the sequence from the new curation. If the missing cytosine is replaced, T3A and T3B will translate to proteins of 228 and 268 amino acids and show 98 and 97% amino acid sequence identity to D. melanogaster. Figure S21. The annotated gene region of Dsim\GD11388 (DsimGSTE11) from Flybase. A. The curated ATG start site is highlighted in yellow, however the end of intron 1 shows a stop codon (TAG) in the annotation but which may be CTG, as in D. melanogaster, in this low coverage genome. B. The figure shows the amino acid alignment of DmelGSTE11, DsimGSTE11 (1) and DsimGSTE11 (2). DsimGSTE11 (1) is the annotated sequence from Flybase and DsimGSTE11 (2) is a new curated sequence. If the stop codon (TAG) at the end of intron 1 is CTG as in D. melanogaster and the translation start site is the yellow highlighted sequence, the proteins show 97% sequence identity to each other. The Ser13 in DmelGSTE11 (shown by arrow) is the catalytic serine and that region is generally conserved for active site topology which suggests that DsimGSTE11 (1) may be an inactive enzyme. Figure S22. The annotated gene region of Dsim\GD24922 (DsimGSTE12) from Flybase. A. Using DmelGSTE12 as template, we found that possibly a thymidine (T) is missing as shown by the underlined gap in exon 2. If this is so, intron 2 also would be translated to the coding sequence. If the thymidine is present a transcription read through would include intron 2 as coding sequence, which then would encode a protein of 223 amino acids. B. The panel shows the amino acid alignment of DmelGSTE12, DsimGSTE12 (1) and DsimGSTE12 (2). DsimGSTE12 (1) is the annotated sequence from Flybase and DsimGSTE12 (2) is a new curated sequence. The new curated sequence of DsimGSTE12 (2) aligns with DmelGSTE12 giving 98% amino acid sequence identity. Figure S23. The ambiguous annotation of Dsec\GM24014 and Dsec\GM24015 (DsecGSTD2). A. The first exon is the annotated coding sequence of Dsec\GM24014 whereas the second exon belongs to Dsec\GM24015 and the small gap of 23 nucleotides between the 2 genes is shown underlined. A missing cytosine of Dsec\GM24014 results in a frame shift which also results in a premature stop codon (TGA). The addition of cytosine to Dsec\GM24014 followed by a read through the gap as well as the exon of Dsec\GM24015 would result in a protein of 215 amino acids. This single exon would encode GSTD2 of D. sechellia. B. The GSTD2 amino acids sequence alignment of D. sechellia and D. melanogaster shows 98% identity. Figure S24. The annotated gene region of Dsec\GM25062 (DsecGSTO4) from Flybase. A. Using DmelGSTO4 as template, we found that intron 2 and part of intron 3 (as underlined) of D. sechellia GSTO4 show very highly conserved gene sequences to DmelGSTO4. However there are two insertions of cytosine, one in exon 2 the other in exon 3 as shown in red. The insertion in exon 2 would lead to a stop codon in intron 2 (TGA as underlined), hence its intron annotation. Moreover, annotated intron 2 contains 1 nucleotide less than DmelGSTO4. This frame shift leads to wrong coding protein. Therefore intron 2 and part of intron 3 were annotated by FlyBase as intron which makes DsecGSTO4 shorter than normal GST Omega class (204 instead of 240 to 250 amino acids). The sequence of intron 3 that is underlined should be part of the protein coding gene, WCERLELLKLQRGEDYNYDESRFPQL. B. The comparison between the conserved block of D. sechellia and D. melanogaster GSTO4. In the annotated intron 2, 7 nucleotides downstream from the stop codon there appears to be a deletion (position shown in panel A as underlined T). If in the exon this one nucleotide deletion would result in a frame-shift. C. The figure shows amino acid alignment of GSTO4 from D. sechellia and D. melanogaster. The empty blocks are the missing sequences of intron 2 and 3. D. The figure shows amino acid alignment of GSTO4 from D. sechellia and D. melanogaster if the inserted cytosines are absent from exons 2 and 3, as in D. melanogaster, and one more nucleotide present in intron 2 (GTGATCTGGATCCCTTCTGGAGCGGCCTGGA CGTCTACGAAAG), the full coding sequence would be present. The proteins would have 97% amino acid sequence identity to each other. This sequence is ambiguous because of the low coverage genome sequencing. Figure S25. The annotated gene region of Dsec\GM21874 (DsecGSTE1) from Flybase. A. A possible ATG start site is located in the intron region. However it includes a TAG stop codon due to the absence of a T, TA(T)G. If the T is present as it is in D. melanogaster, the gene will code for a 206 amino acid protein. B. The GSTE1 amino acid sequence alignment of D. sechellia and D. melanogaster shows 65% identity. Figure S26. The annotated gene region of Dsec\GM21098 (DsecGSTE13) from Flybase. A. In D. melanogaster the sequences from exon 3 to 4 are one continuous exon. There is one adenosine (A) insertion as shown in red which results in a frameshift. Therefore it appears some of the exon region was annotated as intron to keep the rest of the sequence in-frame. B. DsecE13 (1) refers to the sequence reported by Flyabse. DsecE13 (2) refers to the sequence from our new curation. The amino acid sequence alignment of DmelGSTE13 and DsecGSTE13 (1) shows 92% identity. The missing gap is the part of the gene which is annotated as intron. If the extra adenosine is absent as in D. melanogaster, the amino acid identity increases to 98% for 226 amino acids. Figure S27. The annotated gene region of Dvir\GJ14962 (DvirGSTE13) from Flybase. A. We found one adenosine (A) insertion at nucleotide position 419 of exon 4 (shown in red). This insertion makes a frame shift which reads through the stop codon (TAA) shown in red. B. The figure shows the amino acid alignment of DmelGSTE13, DvirGSTE13 (1) and DvirGSTE13 (2). DvirGSTE13 (1) is the annotated sequence from Flybase and DvirGSTE13 (2) is a new curated sequence. The amino acid alignment of DmelGSTE13 and DvirGSTE13 (1) shows 57% identity. If the extra A is absent, as in D. melanogaster, the TAA stop codon shown in red will be in-frame. The two proteins then show 74% amino acid identity to each other. Figure S28. Alternative splicing scheme of DyakGSTD11. A. DyakGSTD11 gene undergoes alternative splicing to generate 2 variants, DyakGSTD11A and DyakGSTD11B. DyakGSTD11Aare from exon 1, 3, 4 and 5. The whole exon 1 and the first 10 nucleotides of exon 3 are proposed to be 5′UTR of DyakGSTD11A whereas the first 17 nucleotides of exon 2 is proposed to be 5′UTR of DyakGSTD11B as shown underlined. DyakGSTD11A and DyakGSTD11B share the sequences in exon 4 and 5. They also share 3′UTR which is the last 94 nucleotides of exon 5 (underline). B. We propose the possible 5′UTR of DyakGSTD11A. This sequence shows nucleotide sequence identity of 83% to D. melanogaster. C. We propose the possible 5′UTR of DyakGSTD11B. This sequence shows nucleotide sequence identity of 100% to D. melanogaster. D. We propose the possible 3′UTR of DyakGSTD11. This sequence shows nucleotide sequence identity of 81% to D. melanogaster. Figure S29. Alternative splicing scheme of DyakGSTZ2. A. DyakGSTZ2 gene undergoes alternative splicing to generate 3 variants, DyakGSTZ2A, DyakGSTZ2B and DyakGSTZ2C. The first 356 nucleotides of exon 1 is proposed to be 5′UTR of DyakGSTZ2B as shown underlined. The first 40 nucleotides of exon 7 and the first 29 nucleotides of exon 6 are proposed to be 5′UTR of DyakGSTZ2C and DyakGSTZ2A, respectively. B. We propose the possible 5′UTR of DyakGSTZ2A. This sequence shows nucleotide sequence identity of 100% to D. melanogaster. C. We propose the possible 5′UTR of DyakGSTZ2B. This sequence shows nucleotide sequence identity of 89% to D. melanogaster. D. We propose the possible 5′UTR of DyakGSTZ2C. This sequence shows nucleotide sequence identity of 95% to D. melanogaster. E. We propose the possible 3′UTR of DyakGSTZ2A. This sequence shows nucleotide sequence identity of 83% to D. melanogaster. F. We propose the possible 3′UTR of DyakGSTZ2B. This sequence shows nucleotide sequence identity of 86% to D. melanogaster. G. We propose the possible 3′UTR of DyakGSTZ2C. This sequence shows nucleotide sequence identity of 85% to D. melanogaster. Figure S30. Alternative splicing scheme of DyakGSTT3. A. DyakGSTT3 gene undergoes alternative splicing to generate 3 variants, DyakGSTT3A, DyakGSTT3B and DyakGSTT3C. The first 235 nucleotides of exon 1 is proposed to be 5′UTR of DyakGSTT3B as shown underlined. Exon 1, 2, 3 and the first 29 nucleotides of exon 4 is proposed to be 5′UTR of DyakGSTT3C. Exon 3 and the first 29 nucleotides of exon 4 is proposed to be 5′UTR of DyakGSTT3A as shown in yellow highlight. All 3 variants share exon 4–6 and the 3′UTR which is the last 165 nucleotides of exon 6 as underlined. B. We propose the possible 5′UTR of DyakGSTT3B. This sequence shows nucleotide sequence identity of 71% to D. melanogaster. C. We propose the possible 5′UTR of DyakGSTT3C. This sequence shows nucleotide sequence identity of 80% to D. melanogaster. D. We propose the possible 5′UTR of DyakGSTT3A. This sequence shows nucleotide sequence identity of 94% to D. melanogaster. E. We propose the possible 3′UTR of DyakGSTT3. This sequence shows nucleotide sequence identity of 72% to D. melanogaster. Figure S31. Alternative splicing scheme of DyakGSTO2. A. Flybase has reported that Dyak\GE21297 and Dyak\GE21298 are GST proteins in Omega class translated from 2 different genes. Our manual curation based on D. melanogaster has shown that these 2 GST proteins could possibly originate from the same gene which has undergone alternative splicing to yield 2 final protein products which are DyakGSTO2A and DyakGSTO2B. The two proteins share sequence in exon 1. The underlined sequence in exon 1 is proposed to be 5′UTR whereas the underlined sequences in exon 2 and 4 are proposed to be 3′UTR of DyakGSTO2B and DyakGSTO2A, respectively. B. We propose the possible 5′UTR of DyakGSTO2. Using D. melanogaster as a template this sequence shows 92% nucleotide sequence identity. C. We propose the possible 3′UTR of DyakGSTO2A. Using D. melanogaster as a template this sequence shows 70% nucleotide sequence identity. D. We propose the possible 3′UTR of DyakGSTO2B. Using D. melanogaster as a template this sequence shows 77% nucleotide sequence identity. Figure S32. The curation for alternative splicing in D. yakuba. Our curation of alternative splicing shows all 11 drosophila had alternatively spliced genes. This figure uses the D. yakuba genes to illustrate this result. The details for this cartoon are given in Table S2. Figure S33. The alternatively spliced GST proteins from the 12 Drosophila species. Each identified alternatively spliced GST from D. melanogaster was used to manually curate the other 11 Drosophila species. This figure shows the amino acid alignment for the spliced GSTs with a matrix table beneath showing the percent identity and directly below the percent similarity for each species GST compared to D. melanogaster's GST template. Figure S34. Genomic organization of the GSTs of the 11 Drosophila species. The genomic maps show the overall synteny of the GST gene family in the Drosophila genus. The genes are shown as arrows which indicate the direction of transcription. The blue color shows genes that FlyBase reports to have GST molecular function but does not specify a D. melanogaster ortholog. The magenta color shows pseudogenes identified by our curation. The purple color shows identified genes needing confirmation due to incomplete genome sequence. Figure S35. Phylogenetic tree analysis The phylogenetic tree analysis used sequences from each GST class from the 12 Drosophila species. Each isoform is designated by their FlyBase symbol followed by the D. melanogaster ortholog name in parenthesis. For example, Dana\GF17052(D1), where D1 refers to the DmelGSTD1 ortholog. In some species there is more than one isoform orthologous to the same D. melanogaster ortholog, so these will be noted as D1-2, D1-3 and so on. GSTT3A/B of D. sechellia and GSTT4 of D. persimilis were not included in this analysis due to incomplete gene sequencing. Table S1. Glutathione transferase orthologs in the 11 Drosophila species. Table S2. Example of alternative splicing curation of non-annotated genes for D. yakuba. (RAR)

创建时间:
2014-10-13
二维码
社区交流群
二维码
科研交流群
商业服务