遇见数据集

Data from: More taxa or more characters revisited: combining data from nuclear protein-encoding genes for phylogenetic analyses of Noctuoidea (Insecta: Lepidoptera)

收藏
DataONE2009-06-15 更新2024-06-27 收录
数据链接:
官方服务:

资源简介:

A central question concerning data collection strategy for molecular phylogenies has been, is it better to increase the number of characters or the number of taxa sampled to improve the robustness of a phylogeny estimate? A recent simulation study concluded that increasing the number of taxa sampled is preferable to increasing the number of nucleotide characters, if taxa are chosen specifically to break up long branches. We explore this hypothesis by using empirical data from noctuoid moths, one of the largest superfamilies of insects. Separate studies of two nuclear genes, elongation factor-1α (EF-1α) and dopa decarboxylase (DDC), have yielded similar gene trees and high concordance with morphological groupings for 49 exemplar species. However, support levels were quite low for nodes deeper than the subfamily level. We tested the effects on phylogenetic signal of (1) increasing the taxon sampling by nearly 60%, to 77 species, and (2) combining data from the two genes in a single analysis. Surprisingly, the increased taxon sampling, although designed to break up long branches, generated greater disagreement between the two gene data sets and decreased support levels for deeper nodes. We appear to have inadvertently introduced new long branches, and breaking these up may require a yet larger taxon sample. Sampling additional characters (combining data) greatly increased the phylogenetic signal. To contrast the potential effect of combining data from independent genes with collection of the same total number of characters from a single gene, we simulated the latter by bootstrap augmentation of the single-gene data sets. Support levels for combined data were at least as high as those for the bootstrap-augmented data set for DDC and were much higher than those for the augmented EF-1α data set. This supports the view that in obtaining additional sequence data to solve a refractory systematic problem, it is prudent to take them from an independent gene.

分子系统发育学的数据收集策略核心争议之一在于:为提升系统发育推断的稳健性,究竟应当增加性状数量,还是扩大采样类群规模?近期一项模拟研究指出,若类群选择的目标为打破长支(long branches),则扩大采样类群数要优于增加核苷酸性状数量。本研究以夜蛾总科(noctuoid moths)——昆虫最大总科之一的实证数据为材料,对该假说进行验证。此前已有两项针对两个核基因的独立研究:延伸因子1α(elongation factor-1α, EF-1α)与多巴脱羧酶(dopa decarboxylase, DDC),二者构建的基因树结果相似,且与49个代表物种的形态学类群划分高度一致,但亚科级以上的深层节点支持率普遍偏低。我们分别测试了两种操作对系统发育信号的影响:(1) 将采样类群数提升近60%至77种,以打破长支;(2) 将两个基因的数据合并后开展联合分析。令人意外的是,尽管本次类群扩增的设计初衷为打破长支,但扩增后两个基因数据集之间的分歧反而增大,深层节点的支持率也有所下降。我们推测这是无意间引入了新的长支,若要打破这些新长支,可能需要进一步扩大类群采样规模。增加性状数量(即合并数据)则显著提升了系统发育信号强度。为对比独立基因合并数据与从单基因中获取相同总性状数的潜在效果差异,我们通过自助法扩增(bootstrap augmentation)单基因数据集的方式模拟了后者。结果显示,合并数据的节点支持率至少不低于DDC数据集的自助法扩增结果,且远高于EF-1α数据集的扩增结果。该结果支持以下观点:为解决棘手的系统发育难题而获取额外序列数据时,从独立基因中取材是更为稳妥的策略。

创建时间:
2009-06-15
二维码
社区交流群
二维码
科研交流群
商业服务