遇见数据集

Phylogenetic analyses of syntenic gene families in visual opsin gene-bearing chromosome regions

收藏
Mendeley Data2024-06-29 更新2024-06-29 收录
官方服务:

资源简介:

Sequence based phylogenetic analyses of 34 vertebrate gene families identified in an analysis of conserved synteny in chromosome regions containing the genes for visual opsins, the G-protein alpha subunit families for transducin (GNAT) and adenylyl cyclase inhibition (GNAI), the oxytocin and vasopressin receptors (OT/VP-R), and the L-type voltage gated calcium channels (CACNA1-L). For each gene family amino acid sequences were predicted from the Ensembl genome browser (http://www.ensembl.org) and used to create sequence alignments and phylogenetic trees. Vertebrate gene families were defined based on Ensembl protein family predictions. Database identifiers, location data, genome assembly information and annotation notes for all identified protein families and sequences are included in 'Supplemental Table 705852.xlsx' (Excel spreadsheet). This spreadsheet also includes informaction on 7 gene families that were discarded from the analyses. Gene families are identified by unique abbreviations based on approved HUGO Gene Nomenclature Committe (HGNC) gene symbols, or known aliases from the NCBI Entrez Gene database. File information: For each gene family an alignment file '...align.fasta', a neighbor joining tree '...NJ.phb' and a phylogenetic maximum likelihood tree '...PhyML.phb' are included. Alignments are included in FASTA format with the extension '.fasta'. This file format can be opened by most sequence analysis applications as well as text editors. Phylogenetic tree files are included in Phylip/Newick format with the extension '.phb'. This file format can be opened by freely available phylogenetic tree viewers such as FigTree (http://tree.bio.ed.ac.uk/software/figtree/) and TreeView (http://darwin.zoology.gla.ac.uk/~rpage/treeviewx/). Corresponding figures for all phylogenetic trees are also included as PDF files. Sequence names/leaf names include species abbreviations (see below) as well as chromosome/linkage group/genomic scaffold numbers, with lowercase letters to distinguish sequences located on the same chromomosome, linkage group or scaffold. For the human sequences the full HGNC gene symbol is included. The species included in these analyses were (abbreviations and common names in parenthesis): Homo sapiens (Hsa, human), Mus musculus (Mmu, mouse), Monodelphis domestica (Mdo, grey short-tailed opossum), Gallus gallus (Gga, chicken), Danio rerio (Dre, zebrafish), Oryzias latipes (Ola, medaka), Gasterosteus aculeatus (Gac, three-spined stickleback), Tetraodon nigroviridis (Tni, green spotted pufferfish), Ciona intestinalis (Cin, tunicate), Ciona savignyi (Csa, tunicate) and Drosophila melanogaster (Dme, fruit fly). In some analyses the following additional species were used: Sarcophilus harrisii (Sha, Tasmanian devil), Taeniopygia guttata (Tgu, zebra finch), Anolis carolinensis (Aca, Carolina anole lizard), Xenopus (Silurana) tropicalis (Xtr, Western clawed frog), Takifugu rubripes (Tru, Japanese pufferfish), Branchiostoma floridae (Bfl, Florida lancelet) and Caenorhabditis elegans (Cel, nematode). The following vertebrate gene families are included in this file set: ATP2B: ATPase, Ca++ transporting, plama membrane B4GALNT: Beta-1,4-N-acetyl-galactosaminyl transferase CACNA2D: Calcium channel, voltage-dependent, alpha 2/delta subunit CAMK1: Calcium/calmodulin dependent protein kinase CDK: Cyclin-dependent kinase, members 16, 17 and 18 CELSR: Cadherin, EGF LAG seven-pass G-type receptor (flamingo homolog, Drosophila) CNTN: Contactin precursor COPG: Coatomer protein complex, subunit gamma ERC: ELKS/RAB6-interacting/CAST family FLN: Filamin GXYLT: Glucoside xylosyltransferase IKBKE: Kinase epsilon and TANK-binding kinase IQSEC: IQ motif and Sec7 domain containing KDM: Lysine specific demethylase 5 KLHDC: Kelch domain containing 8 L1CAM: L1 cell adehesion molecule LRRN: Leucine rich repeat neuronal MAGI: Membrane associated guanylate kinase, WW and PDZ domain containing PHTF: Putative homeodomain transcription factor PLG: Plaminogen ortholog PLXNA: Plexin A PPM1: Protein phosphatase, Mg2+/Mn2+ dependent PRICKLE: Prickle homolog PTPN: Protein tyrosine phosphatase, non-receptor type RBM: RNA binding motif protein RSBN: Round spermatid basic protein SEMA3: Sema domain, immunoglobulin domain (Ig), short basic domain, secreted, (semaphorin) SRGAP: SLIT-ROBO Rho GTPase activating protein SYP: Synaptophysin TIMM: Translocase of inner mitochondrial membrane 17 TWF: Twinfilin UBA: Ubiquitin-like modifier activating enzyme, members 1 and 7 USP: Ubiquitin specific peptidase, members 4, 11, 15 and 19 WNK: WNK lysine deficient protein kinase Method details: Alignments were created using the ClustalW sequence alignment algorithm with the following settings: Gonnet weight matrix, gap opening penalty 10.0 and gap extension penalty 0.20. Phylogenetic analyses were carried out based on the included alignments using bootstrap-supported neighbor joining (NJ) as well as phylogenetic maximum likelihood (PhyML) methods supported by approximate likelihood ratio tests (aLRT). Phylogenetic trees are rooted with identified Drosophila melanogaster (fruit fly) sequences, if possible. Alternatively some phylogenetic trees are rooted with other identified invertebrate sequences (see Supplemental Table 1). The B4GALNT, PLG, PTPN, RBM, SEMA3 and USP trees are presented as mindpoint-rooted trees in the figures (PDF), however the phylogenetic tree files (.phb) are unrooted. NJ trees were made using standard settings in ClustalX 2.0.12 (http://www.clustal.org/clustal2/), supported by a non-parametric bootstrap analysis with 1000 replicates. PhyML trees were made using the PhyML3.0 algorithm (http://www.atgc-montpellier.fr/phyml/‎) with the following settings: amino acid frequencies (equilibrium frequencies), proportion of invariable sites (with optimised p-invar) and gamma shape parameters were estimated from the alignments, the number of substitution rate categories was set to 8, BIONJ was chosen to create the starting tree, both NNI and SPR tree optimization methods were considered and both tree topology and branch length optimization were chosen. The amino acid substitution model was chosen based on ProtTest3.2 (http://code.google.com/p/prottest3/) results. The JTT model was applied for all gene families except B4GALNT, CACNA2D, COL, L1CAM, PLG, PPP, QSOX and UBA where the WAG model was chosen, and RPL and TWF where the LG model was chosen. PhyML trees are supported by approximate likelihood ratio tests (aLRT) with SH-like branch upports applied through PhyML. For the CAMK and GXYLT gene families the PhyML trees were repeated (same settings) using a non-parametric bootstrap analysis with 100 replicates rather that aLRT in PhyML. These trees did not improve on the aLRT-supported tree topologies.

本数据集包含针对34个脊椎动物基因家族的基于序列的系统发育分析结果,该分析基于对包含视觉视蛋白、转导蛋白G蛋白α亚基家族(GNAT)、腺苷酸环化酶抑制型G蛋白α亚基家族(GNAI)、催产素与血管加压素受体(OT/VP-R)以及L型电压门控钙通道(CACNA1-L)的染色体区域进行保守同线性分析所鉴定得到的基因家族。 所有基因家族的氨基酸序列均从Ensembl基因组浏览器(Ensembl Genome Browser,http://www.ensembl.org)预测得到,用于构建序列比对文件与系统发育树。脊椎动物基因家族的定义基于Ensembl蛋白家族预测结果。所有鉴定得到的蛋白家族与序列的数据库标识符、位置信息、基因组组装信息及注释说明均收录于Supplemental Table 705852.xlsx(Excel电子表格)中,该表格同时包含分析过程中被舍弃的7个基因家族的相关信息。基因家族通过基于人类基因命名委员会(HUGO Gene Nomenclature Committee, HGNC)批准的基因符号或NCBI Entrez基因数据库中的已知别名所对应的唯一缩写进行标识。 本数据集的文件组成如下:每个基因家族对应一个FASTA格式的序列比对文件(后缀为.align.fasta)、一个邻接法系统发育树文件(后缀为.NJ.phb)以及一个最大似然法系统发育树文件(后缀为.PhyML.phb)。其中,序列比对文件采用FASTA格式,可通过绝大多数序列分析软件及文本编辑器打开;系统发育树文件采用Phylip/Newick格式(后缀为.phb),可通过FigTree(http://tree.bio.ed.ac.uk/software/figtree/)与TreeView(http://darwin.zoology.gla.ac.uk/~rpage/treeviewx/)等免费系统发育树查看软件打开。所有系统发育树对应的可视化结果均以PDF文件形式提供。 序列名称/叶节点名称包含物种缩写、染色体/连锁群/基因组支架编号,若同一条染色体、连锁群或支架上存在多个序列,则通过小写字母进行区分。人类序列会附带完整的HGNC基因符号。 本分析涉及的物种(缩写与俗名附于括号内)包括:智人(Homo sapiens, Hsa, 人类)、小家鼠(Mus musculus, Mmu, 小鼠)、灰短尾负鼠(Monodelphis domestica, Mdo, 灰短尾负鼠)、原鸡(Gallus gallus, Gga, 鸡)、斑马鱼(Danio rerio, Dre, 斑马鱼)、青鳉(Oryzias latipes, Ola, 青鳉)、三刺鱼(Gasterosteus aculeatus, Gac, 三刺棘鱼)、绿斑河豚(Tetraodon nigroviridis, Tni, 绿斑河豚)、柄海鞘(Ciona intestinalis, Cin, 被囊动物)、萨氏海鞘(Ciona savignyi, Csa, 被囊动物)以及黑腹果蝇(Drosophila melanogaster, Dme, 果蝇)。部分额外分析使用的物种包括:袋獾(Sarcophilus harrisii, Sha, 塔斯马尼亚恶魔)、斑胸草雀(Taeniopygia guttata, Tgu, 斑胸草雀)、卡罗莱纳安乐蜥(Anolis carolinensis, Aca, 卡罗莱纳安乐蜥)、西方爪蟾(Xenopus (Silurana) tropicalis, Xtr, 西方爪蟾)、红鳍东方鲀(Takifugu rubripes, Tru, 日本河豚)、佛罗里达文昌鱼(Branchiostoma floridae, Bfl, 佛罗里达文昌鱼)以及秀丽隐杆线虫(Caenorhabditis elegans, Cel, 线虫)。 本数据集包含的脊椎动物基因家族列表如下: - ATP2B:质膜型钙转运ATP酶 - B4GALNT:β-1,4-N-乙酰半乳糖胺基转移酶 - CACNA2D:电压依赖性钙通道α2/δ亚基 - CAMK1:钙/钙调蛋白依赖性蛋白激酶 - CDK:细胞周期蛋白依赖性激酶(成员16、17和18) - CELSR:钙黏蛋白EGF LAG七次跨膜G型受体(果蝇flamingo同源物) - CNTN:接触蛋白前体 - COPG:衣被蛋白复合物γ亚基 - ERC:ELKS/RAB6相互作用/CAST家族 - FLN:细丝蛋白 - GXYLT:葡萄糖苷木糖基转移酶 - IKBKE:激酶ε与TANK结合激酶 - IQSEC:含IQ基序与Sec7结构域蛋白 - KDM:赖氨酸特异性去甲基化酶5 - KLHDC:含Kelch结构域蛋白8 - L1CAM:L1细胞黏附分子 - LRRN:亮氨酸重复序列神经元蛋白 - MAGI:含WW和PDZ结构域的膜相关鸟苷酸激酶 - PHTF:推定同源域转录因子 - PLG:纤溶酶原同源物 - PLXNA:丛蛋白A - PPM1:Mg²+/Mn²+依赖性蛋白磷酸酶 - PRICKLE:Prickle同源物 - PTPN:非受体型蛋白酪氨酸磷酸酶 - RBM:RNA结合基序蛋白 - RSBN:圆形精子碱性蛋白 - SEMA3:含Sema结构域、免疫球蛋白结构域(Ig)、短碱性结构域的分泌型信号素 - SRGAP:SLIT-ROBO Rho GTP酶激活蛋白 - SYP:突触泡蛋白 - TIMM:线粒体内膜转位酶17 - TWF:丝束蛋白 - UBA:泛素样修饰物激活酶(成员1和7) - USP:泛素特异性肽酶(成员4、11、15和19) - WNK:赖氨酸缺陷型WNK蛋白激酶 方法细节: 序列比对采用ClustalW算法完成,参数设置为:Gonnet权重矩阵、空位开放罚分10.0、空位延伸罚分0.20。系统发育分析基于上述比对序列,分别采用bootstrap支持的邻接法(Neighbor Joining, NJ)以及近似似然比检验(approximate likelihood ratio tests, aLRT)支持的最大似然法(PhyML)进行。若可获取,系统发育树以鉴定出的黑腹果蝇序列作为外类群定根;部分树则以其他鉴定出的无脊椎动物序列定根(详见Supplemental Table 1)。B4GALNT、PLG、PTPN、RBM、SEMA3及USP家族的系统发育树在PDF可视化图中采用中点定根,但对应的.phb树文件未进行定根。邻接法树通过ClustalX 2.0.12(http://www.clustal.org/clustal2/)的默认参数构建,并采用1000次重复的非参数bootstrap分析进行支持验证。最大似然法树采用PhyML 3.0算法(http://www.atgc-montpellier.fr/phyml/)构建,参数设置如下:从比对序列中估算氨基酸频率(平衡频率)、不变位点比例(优化p-invar参数)以及伽马形状参数;替换速率分类数设为8;以BIONJ算法生成初始树;同时采用NNI和SPR两种树拓扑优化方法,并同时优化树拓扑结构与分支长度。氨基酸替换模型根据ProtTest 3.2(http://code.google.com/p/prottest3/)的分析结果选取:除B4GALNT、CACNA2D、COL、L1CAM、PLG、PPP、QSOX及UBA家族采用WAG模型,RPL与TWF家族采用LG模型外,其余所有基因家族均采用JTT模型。PhyML树通过近似似然比检验(aLRT)结合SH-like分支支持值进行验证,该验证由PhyML内置实现。针对CAMK1与GXYLT基因家族,我们采用100次重复的非参数bootstrap分析替代PhyML内置的aLRT检验,重复构建PhyML树(参数一致);但最终树拓扑结构并未优于aLRT支持的树拓扑。

创建时间:
2023-06-28
二维码
社区交流群
二维码
科研交流群
商业服务