Some characteristics of the 12 LEAP classes.
收藏资源简介:
aMeaning of the regular expression syntax used for motifs: «.» = any amino acid; X{n, } = at least n times X; X{n,m} = n to m times X; [XY] = X or Y; [∧XY] = neither X nor Y; (XY) = X followed by Y; X? = X present or not; XY$ = XY at the end; (M1)|(M2) = motif M1 or motif M2 or both.bNumber of sequences in LEAPdb using the motif indicated.cAmino acid sequences length range of LEAP classes in LEAPdb.dConsensus sequences of the LEAP classes obtained using Multalin [74]: alignment of all sequences of each LEAP class was performed with a low consensus value (CV) = 35% and a high consensus value = 60% (i.e., above the «twilight zone» [31]) with a PAM matrix (since sequences of each LEAPs class are either distant or not). Gap penalties values (gap open penalty = 2/gap extension penalty = 0/no gap penalty for extremities) were chosen in order to have not stringent conditions for the alignments, thus introducing numerous gaps (see the gaps percentage). This «local - global alignment» of each LEAP class sequences leads to a consensus sequence for each LEAP class, revealing a high level of similarity between those sequences (also much above the «twilight zone»), especially in the case of LEAP classes 3, 5, 7, 10, 11 and 12.eMotif class 4: STTAPGHY|HKTGTTTS|GGGGIGTG|HS[DR]N?K$|DVE$|LH(TRASHEES)?$|C?TGH$|DKLPGQH$|QQN(KTGCD)?$| RGD$|KEGY$|GHRPQI$|GHNN$|SFKS$|GTHKGL$|SSRDNY$|GQSK$|HRDV$|NDL$.fMotif class 6: [∧LNP][∧G][ADEGILMQRSTVY] [AEKQRSTY].[KR][AT].[ADENT][∧DP][EGIKLMQST].{1,67}[∧DER][∧AS]K[AD][∧IL][∧N].[∧E]?.{1,6}G?
a. 用于基序(motif)的正则表达式语法含义:«.» 代表任意氨基酸;X{n, } 表示字符X至少重复n次;X{n,m} 表示字符X重复n至m次;[XY] 表示X或Y;[∧XY] 代表既非X也非Y;(XY) 表示X后紧跟Y;X? 表示X存在或不存在;XY$ 表示XY位于序列末端;(M1)|(M2) 表示基序M1、M2或二者兼具。 b. LEAPdb中使用该指定基序的序列条数。 c. LEAPdb内LEAP类的氨基酸序列长度范围。 d. 采用Multalin [74] 工具得到的LEAP类共有序列:对每个LEAP类的全部序列进行多重序列比对,设置低一致性阈值(consensus value, CV)=35%、高一致性阈值=60%(即高于"暮光区(twilight zone)"[31]),并选用PAM矩阵(因各LEAP类序列或亲缘关系较远,或并无显著亲缘性)。间隙罚分参数设定为:开放间隙罚分=2、延伸间隙罚分=0、末端无间隙罚分,以采用较为宽松的比对条件,从而引入大量间隙(详见间隙百分比说明)。对每个LEAP类序列执行此"局部-全局比对"后,可得到对应LEAP类的共有序列,结果显示各序列间相似度极高(亦远高于暮光区阈值),尤其在LEAP类3、5、7、10、11及12中表现尤为显著。 e. 基序4类:STTAPGHY|HKTGTTTS|GGGGIGTG|HS[DR]N?K$|DVE$|LH(TRASHEES)?$|C?TGH$|DKLPGQH$|QQN(KTGCD)?$| RGD$|KEGY$|GHRPQI$|GHNN$|SFKS$|GTHKGL$|SSRDNY$|GQSK$|HRDV$|NDL$。 f. 基序6类:[∧LNP][∧G][ADEGILMQRSTVY] [AEKQRSTY].[KR][AT].[ADENT][∧DP][EGIKLMQST].{1,67}[∧DER][∧AS]K[AD][∧IL][∧N].[∧E]?.{1,6}G?



