Code and data for: Evolutionary assembly of the plant terrestrialization toolkit from protein domains
收藏资源简介:
Evolutionary assembly of the plant terrestrialization toolkit from protein domains We want to trace the evolutionary predisposition the green lineage had at the time of terrestrialization. Terrestrialization is a complete change in environment from water, and these acted like new and/or intensified stress factors on the green lineage. Thus, in order to trace the evolutionary predisposition which led to successful overcoming of challenges of a new environment, we carefully annotate, dissect and analyse the stress proteome of the extant Green lineage. Key Concepts Already existing Concept: Orthology New/Redefined Concept Embryophytic Domain: Protein domain which has an Embryophytic ancestor, that is, present in at least one Bryophyte species and at least one Tracheophyte species Latent Genetic Potential (LGP): Key (functional) Embryophytic protein domains having Last Common Ancestor (LCA) prior to the LCA of protein with corresponding function Data: (Figure 1 A) We take 6 species from Cholorophyte algae lineage, 7 from Streptophyte algae lineage, 5 from Bryophyte lineage and 14 from Tracheophyte lineage. 5 species from Cyanobacteria were taken as an outgroup lineage. Stress-annotation (Figure 1B) First 57,796 orthogroups were obtained from 37 proteome. This is computed by using OrthoFinder2. Next we need to annotate stress-related proteins across our dataset. Following are existing ways of annotation: Extensive experiments have been done on A.thaliana and P.patens in the green lineage. Thus, we stress-annotated 37 proteomes using TAIR10 and PEATmoss respectively. From this method, we stress-annotated 4,902 orthogroups Eggnogmapper is a known tool with an extensive database used for gene ontology purposes. We stress-annotated 17,349 orthogroups. In order to avoid experimental bias and tool-based randomness we overlap the two methods to have our final stress-annotation. We finally filtered 2,475 stress-annotated orthogroups from 57,796 orthogroups. Overview of Stress-annotation (Figure 1C, 1D) To have an overview of stress-annotation, overlaps of stress-annotated protein domains, proteins and orthogroups across our dataset lineages and various response to stresses are shown. Distribution of Stress-annotation (Figure 2) Using the overlap annotation approach explained earlier, the stress-annotated orthogroups and corresponding protein domains are in Figure 2A. A standard pattern is observed in both, that is, the unique number of stress orthogroups and protein domains increase from Cyanobacteria to Tracheophytes. The number of proteins and average number of protein domains also increases similarly. Top 10 bursts of protein domains from one lineage to the next is shown in Figure 2B. Changes in stress-annotation with respect to Protein Domains (Figure 3) Figure 3 is a 4-dimensional plot with the following parameters: Species(37), Protein Domains (100), Number of Orthgogroups with Protein Domain (size of circle), Number of Proteins with Protein Domain (Color of circle). The plot is sorted from top to bottom based on the number of orthogroups, and the top 100 protein domains are chosen for the plot. This plot is used to express an overview of the most significant occurances of sub- and neo-functionalizations. Here, each orthogroup is considered to be a protein family. 2 protein families can have an overlapping number of functions. That is why there are more than 1 orthogroup which have the same protein domain. Assembling LGP from protein domains (Figure 4) - refer to Key Concepts to understand LGP In Figure 1A, we show 2 categories (x/y) of orthogroups at each node (a,b,c,d,e,f). The number y for example at node b indicates the number of orthogroups (4) in Tracheophyta+Bryophyta+Zygnemaotphyceae that have LGP (or key Embryophytic protein domains) in Charophyceae. The number x at node b indicates the number of orthogroups (131) in Tracheophyta+Bryophyta+Zygnemaotphyceae that have LGP in all the rest of the lineages (Charophyceae+Klebsormidiophyceae+Chlorokybophyceae+Mesostigmatophyceae+Chlorophytes) in the figure. Since we are concerned about the LGP for Land Plants (Embryophytes), we look at node a. Next, we functionally annotate 96 orthogroups. 50 annotations that occur the most number of times is shown in Figure 4B. In Figure 4C, we can see in which species the key Embryophytic Domains are present whose proteins and protein families are only seen in Embryophytes. Thus, from the final figure we can trace the LGP present at the time of terrestrialization in the LCA of Land Plants. Database files: These are intermediate files used in code for different figures. Following is the link to access them:https://data.mendeley.com/datasets/mnrn7j7hrw/draft?a=b981b40f-01a8-48ff-9d6a-151f6223810c [OR]https://owncloud.gwdg.de/index.php/s/dH3Y4MAHSfbmhrA
# 基于蛋白质结构域的植物陆生化工具库演化组装 本研究旨在追溯绿色植物谱系在陆生化(登陆)时期所具备的演化预适应能力。 陆生化是生物生存环境从水生到陆生的彻底转变,这一转变对绿色植物谱系构成了全新的、或强度升级的胁迫因子。因此,为了追溯助力绿色植物成功克服全新环境挑战的演化预适应机制,本研究对现存绿色植物谱系的胁迫蛋白质组进行了精准注释、拆解与分析。 ## 核心概念 ### 已有概念 直系同源(Orthology) ### 新定义/重新定义概念 胚植物结构域(Embryophytic Domain):指具有胚植物祖先起源的蛋白质结构域,即至少在一种苔藓植物(Bryophyte)物种与一种维管植物(Tracheophyte)物种中均存在的蛋白质结构域。 潜在遗传潜能(Latent Genetic Potential, LGP):指具有对应功能的蛋白质的最近共同祖先(Last Common Ancestor, LCA)出现之前就已存在的关键功能型胚植物蛋白质结构域。 ## 数据集(图1A) 本研究选取了6种绿藻门(Chlorophyte algae)物种、7种链型藻(Streptophyte algae)物种、5种苔藓植物(Bryophyte)物种以及14种维管植物(Tracheophyte)物种,同时选取5种蓝细菌(Cyanobacteria)物种作为外类群。 ## 胁迫注释(图1B) 首先从37份蛋白质组中鉴定得到57796个同源基因家族(orthogroup),该步骤通过OrthoFinder2工具完成。 接下来需对本数据集内的胁迫相关蛋白质进行注释,现有注释方法如下: 1. 针对绿色植物谱系中的拟南芥(Arabidopsis thaliana, A.thaliana)与小立碗藓(Physcomitrium patens, P.patens)已有大量研究基础,因此本研究分别采用TAIR10与PEATmoss数据库对37份蛋白质组进行胁迫注释,共注释得到4902个同源基因家族。 2. Eggnogmapper是一款拥有大规模注释数据库的通用基因本体(Gene Ontology, GO)注释工具,本研究通过该工具共注释得到17349个同源基因家族。 为避免实验偏差与工具引入的随机性,本研究对两种方法的注释结果取交集,最终从57796个同源基因家族中筛选得到2475个经胁迫注释的同源基因家族。 ## 胁迫注释概况(图1C、1D) 为直观呈现胁迫注释的整体特征,本研究展示了数据集各谱系中经胁迫注释的蛋白质结构域、蛋白质及同源基因家族的重叠情况,以及各类胁迫响应的分布特征。 ## 胁迫注释分布(图2) 采用前述交集注释方法得到的经胁迫注释的同源基因家族及其对应蛋白质结构域分布如图2A所示。两类数据均呈现一致的分布规律:从蓝细菌到维管植物,独特的胁迫相关同源基因家族与蛋白质结构域数量逐步递增;蛋白质总数与平均蛋白质结构域数量亦呈现类似的增长趋势。图2B展示了各谱系间蛋白质结构域扩张速率排名前十的事件。 ## 基于蛋白质结构域的胁迫注释变化(图3) 图3为四维散点图,其参数设置如下:横轴与纵轴维度分别为37个物种与100个蛋白质结构域,圆点大小代表携带该结构域的同源基因家族数量,圆点颜色代表携带该结构域的蛋白质数量。该图表按同源基因家族数量从多到少排序,仅展示排名前100的蛋白质结构域。 该图表用于直观展示亚功能化与新功能化的典型发生事件。本研究中每个同源基因家族对应一个蛋白质家族,不同蛋白质家族可拥有重叠的功能,因此同一蛋白质结构域可对应多个同源基因家族。 ## 基于蛋白质结构域的潜在遗传潜能(LGP)组装(图4) 请参考前文核心概念部分理解LGP的定义。 图1A展示了各演化节点(a、b、c、d、e、f)处两类同源基因家族(x与y)的分布情况。以节点b为例,数值y代表在双星藻纲(Zygnematophyceae)、苔藓植物与维管植物中,且在轮藻纲(Charophyceae)中携带LGP(即关键胚植物结构域)的同源基因家族数量(共4个);数值x则代表在双星藻纲、苔藓植物与维管植物中,且在其余所有谱系(轮藻纲、链孢藻纲(Klebsormidiophyceae)、绿球藻纲(Chlorokybophyceae)、中球藻纲(Mesostigmatophyceae)以及绿藻门)中携带LGP的同源基因家族数量(共131个)。 由于本研究聚焦于陆生植物(胚植物)的LGP,因此重点关注节点a。随后本研究对96个同源基因家族进行了功能注释,出现频次最高的50个功能注释结果如图4B所示。图4C展示了仅在胚植物中存在的关键胚植物结构域对应的蛋白质及蛋白质家族的物种分布情况。 综上,通过最终的可视化结果,本研究可追溯至陆生植物最近共同祖先在登陆时期所携带的潜在遗传潜能。 ## 数据库文件 本研究用于生成各图表的中间代码文件可通过以下链接获取:https://data.mendeley.com/datasets/mnrn7j7hrw/draft?a=b981b40f-01a8-48ff-9d6a-151f6223810c 或 https://owncloud.gwdg.de/index.php/s/dH3Y4MAHSfbmhrA



