File S1 - The Venom Gland Transcriptome of <i>Latrodectus tredecimguttatus</i> Revealed by Deep Sequencing and cDNA Library Analysis
收藏资源简介:
Figure S1, The distribution of EST sequences in different GO categories. Figure S2, Distribution of the length of identified transcripts. Figure S3, Scatter plots of RPKM distribution of genes in the KEGG classes. Figure S4, The mapping of identified transcripts/proteins (marked as red) in the spliceosome pathway of KEGG database. Figure S5, The mapping of identified transcripts /proteins (marked as red) in the pathway of protein process in endoplasmic reticulum in KEGG database. Figure S6, Sequence characteristics of members in orphan families. Sequence characteristics of members in orphan families. The potential toxins, which haven’t homologue’s function annotations, are classified into orphan families including two groups: one comprises toxins predicted from Cys patterns, and the other is based on sequence homology with known toxins containing domains. Within Cys patterns, the char “#” represents any three amino acids other than Cys. For other toxins, the domain architectures were predicted by the SMART and Pfam servers [59,60]. The character “-F” appended to protein ID numbers indicates that these sequences are fragments but not full-length proteins. The abbreviations of domain names are as follow: EGF (SMART ID: SM00181); KU (SMART ID: SM00131), glyco_hydro_56 (Pfam ID: PF01630); crust_neurohorm (Pfam ID: PF01147). Figure S7, The abundance of toxin families in different functional categories. Bars represent toxin families clustered based on their functional characteristics. The sum of RPKM values for each class and category are labeled. Neurotoxins including the ANK superfamily, the SCP family and the lycotoxin family; Assistant toxins including theriditoxin family; Proteases including ctenitoxin family; Function unknown toxins including scorpion toxin like family and the orphan family. Figure S8, Phylogenomic trees for trypsin, scorpion toxin-like, lycotoxin, ctenitoxin, SCP family. Phylogenomic trees of trypsin, scorpion toxin-like, lycotoxin, ctenitoxin and SCP families. A. Ctenitoxin family; B. Trypsin family; C. Scorpion toxin-like family; D. SCP family; E. Lycotoxin family. The members of family and their homologues from other spiders are colored as blue and red on branches. For spider species that have transcriptomic data were highlighted by a green line. Figure S9, Phylogenetic tree of ANK superfamily toxins and their homologues from other 45 species. Phylogenetic tree of ANK superfamily toxins and their homologues from other 45 species. Color code: pink for α-LTX-Lt1a family1; blue for α-LTX-Lt1a family2; green for δ-LIT-Lt1a family; red for α-LIT-Lt1a family; brown for ANK family. All phylogeny analyses are performed with MEGA 5.2 using Maximum Likelihood algorithm and 1000 bootstrap tests. The numbers on the branches are the supporting percentages of 1000 bootstrap tests. Table S1, RPKM distribution in the top ten of three GO namespaces. Table S2, RPKM statistics of Ion channel in Latrodectus tredecimguttatus. Table S3, The statistics of RPKM in KEGG pathway superclass. Table S4, The RPKM list of sub-classes of the “Genetic information processing” category in KEGG database. Table S5, List of toxins identified by sequence analyses. Table S6, Known ion channel toxins in five venomous species. Table S7, Full names/abbreviations’ and taxonomic classification of 18 species in phylogenetic analysis. Table S8, Full names/abbreviations and taxonomic classification of 54 arthropod species. Table S9, Full names/abbreviations and taxonomic classification of species shown in Figure S8. (PDF)
补充图S1:EST(表达序列标签,Expressed Sequence Tag)序列在不同GO(基因本体,Gene Ontology)功能类别中的分布。 补充图S2:已鉴定转录本的长度分布。 补充图S3:基因RPKM(每百万映射读取的每千碱基片段数,Reads Per Kilobase per Million mapped reads, RPKM)分布在KEGG(京都基因与基因组百科全书,Kyoto Encyclopedia of Genes and Genomes)类别下的散点图。 补充图S4:已鉴定转录本/蛋白质(标记为红色)在KEGG数据库剪接体通路中的定位。 补充图S5:已鉴定转录本/蛋白质(标记为红色)在KEGG数据库内质网蛋白质加工通路中的定位。 补充图S6:孤儿家族成员的序列特征。孤儿家族成员的序列特征。未获得功能注释同源物的潜在毒素被归类为孤儿家族,分为两类:一类为基于半胱氨酸模式预测的毒素,另一类为与含结构域的已知毒素具有序列同源性的毒素。在半胱氨酸模式中,字符"#"代表半胱氨酸以外的任意三种氨基酸。对于其余毒素,其结构域架构通过SMART(简单模块化结构研究工具,Simple Modular Architecture Research Tool)和Pfam(蛋白质家族数据库,Protein family)服务器[59,60]进行预测。蛋白质ID后附加的"-F"表示这些序列为片段而非全长蛋白质。结构域名称缩写如下:EGF(SMART编号:SM00181);KU(SMART编号:SM00131);glyco_hydro_56(Pfam编号:PF01630);crust_neurohorm(Pfam编号:PF01147)。 补充图S7:不同功能类别中毒素家族的丰度。柱状图代表基于功能特征聚类的毒素家族。标注了每个类别下的RPKM值总和。神经毒素包含ANK(锚蛋白,Ankyrin)超家族、SCP家族以及狼蛛毒素家族;辅助毒素包括Theriditoxin家族;蛋白酶包括ctenitoxin家族;功能未知的毒素包括类蝎毒素家族以及孤儿家族。 补充图S8:胰蛋白酶、类蝎毒素、狼蛛毒素、ctenitoxin以及SCP家族的系统发育基因组树。胰蛋白酶、类蝎毒素、狼蛛毒素、ctenitoxin和SCP家族的系统发育基因组树。A. ctenitoxin家族;B. 胰蛋白酶家族;C. 类蝎毒素家族;D. SCP家族;E. 狼蛛毒素家族。分支上,该家族成员及其来自其他蜘蛛的同源物分别以蓝色和红色标注。带有转录组数据的蜘蛛物种以绿色线条高亮显示。 补充图S9:ANK超家族毒素及其来自其他45个物种的同源物的系统发育基因组树。ANK超家族毒素及其来自其他45个物种的同源物的系统发育基因组树。颜色代码:粉色代表α-LTX-Lt1a家族1;蓝色代表α-LTX-Lt1a家族2;绿色代表δ-LIT-Lt1a家族;红色代表α-LIT-Lt1a家族;棕色代表ANK家族。所有系统发育分析均使用MEGA 5.2软件,采用最大似然算法并进行1000次自展检验。分支上的数字为1000次自展检验的支持率百分比。 补充表S1:三个GO基因本体命名空间的前十条RPKM分布。 补充表S2:红斑寇蛛(Latrodectus tredecimguttatus)离子通道的RPKM统计数据。 补充表S3:KEGG通路超类别的RPKM统计数据。 补充表S4:KEGG数据库中"遗传信息处理"类别下子类别的RPKM列表。 补充表S5:通过序列分析鉴定的毒素列表。 补充表S6:五种有毒物种中的已知离子通道毒素。 补充表S7:系统发育分析中涉及的18个物种的全称/缩写及其分类学信息。 补充表S8:54个节肢动物物种的全称/缩写及其分类学信息。 补充表S9:图S8中展示的物种的全称/缩写及其分类学信息。 (PDF格式)



