遇见数据集

Data from: Serine codon-usage bias in deep phylogenomics: pancrustacean relationships as a case study

收藏
DataONE2012-10-12 更新2024-06-27 收录
数据链接:
官方服务:

资源简介:

Phylogenomic analyses of ancient relationships are usually performed using amino acid data, but it is unclear whether amino acids or nucleotides should be preferred. With the 2-fold aim of addressing this problem and clarifying pancrustacean relationships, we explored the signals in the 62 protein-coding genes carefully assembled by Regier et al. in 2010. With reference to the pancrustaceans, this data set infers a highly supported nucleotide tree that is substantially different to the corresponding, but poorly supported, amino acid one. We show that the discrepancy between the nucleotide-based and the amino acids-based trees is caused by substitutions within synonymous codon families (especially those of serine—TCN and AGY). We show that different arthropod lineages are differentially biased in their usage of serine, arginine, and leucine synonymous codons, and that the serine bias is correlated with the topology derived from the nucleotides, but not the amino acids. We suggest that a parallel, partially compositionally driven, synonymous codon-usage bias affects the nucleotide topology. As substitutions between serine codon families can proceed through threonine or cysteine intermediates, amino acid data sets might also be affected by the serine codon-usage bias. We suggest that a Dayhoff recoding strategy would partially ameliorate the effects of such bias. Although amino acids provide an alternative hypothesis of pancrustacean relationships, neither the nucleotides nor the amino acids version of this data set seems to bring enough genuine phylogenetic information to robustly resolve the relationships within group, which should still be considered unresolved.

针对深层演化关系的系统发育基因组学分析通常采用氨基酸数据,但学界尚未明确氨基酸与核苷酸序列何者更具优势。为解决这一问题并阐明泛甲壳动物(pancrustaceans)的演化关系,我们针对Regier等人2010年精心组装的62个蛋白编码基因序列展开了信号分析。针对泛甲壳动物类群,本数据集基于核苷酸序列构建的系统发育树具有极高的支持度,但其与基于氨基酸序列构建的对应拓扑结构差异显著,且后者的支持度极低。研究发现,核苷酸与氨基酸序列推导的系统发育树之间的分歧,源于同义密码子家族内的替换(尤其是丝氨酸密码子TCN和AGY家族)。我们证实,不同节肢动物支系在丝氨酸、精氨酸与亮氨酸的同义密码子使用上存在偏倚差异,且这种丝氨酸密码子使用偏倚与基于核苷酸序列推导的拓扑结构显著相关,却与氨基酸序列推导的拓扑结构无关。我们推测,一种由碱基组成偏倚部分驱动的平行同义密码子使用偏倚,影响了核苷酸序列推导的拓扑结构。由于丝氨酸密码子家族间的替换可通过苏氨酸或半胱氨酸作为中间态实现,氨基酸数据集也可能受到丝氨酸密码子使用偏倚的影响。我们建议采用Dayhoff重编码(Dayhoff recoding)策略,可部分缓解此类偏倚带来的影响。尽管氨基酸序列为泛甲壳动物的演化关系提供了另一套假说,但本数据集的核苷酸与氨基酸版本均未提供足够可靠的系统发育信息,无法稳健解析该类群内部的演化关系,相关问题仍应视为未解决状态。

创建时间:
2012-10-12
二维码
社区交流群
二维码
科研交流群
商业服务