Redefining possible: Combining phylogenomic and supersparse data in frogs
收藏资源简介:
The data available for reconstructing molecular phylogenies have become wildly disparate. Phylogenomic studies can generate data for thousands of genetic markers for dozens of species, but for hundreds of other taxa, data may be available from only a few genes. Can these two types of data be integrated to combine the advantages of both, addressing the relationships of hundreds of species with thousands of genes? Here we show that this is possible, using data from frogs. We generated a phylogenomic dataset for 138 ingroup species and 3,784 nuclear markers (ultraconserved elements, UCEs), including new UCE data from 70 species. We also assembled a supermatrix dataset, including data from 97% of frog genera (441 total), with 1â307 genes per taxon. We then produced a combined phylogenomic-supermatrix dataset (a âgigamatrixâ) containing 441 ingroup taxa and 4,091 markers, but with 86% missing data overall. Likelihood analysis of the gigamatrix yielded a generally well-supported tree among fa..., ,
可用于重建分子系统发育的数据集已呈现出显著的异质性。系统基因组学研究可针对数十个物种生成数千个遗传标记的数据,但针对数百个其他类群,往往仅能获得少数几个基因的序列数据。能否整合这两类数据以兼顾二者优势,从而解析涵盖数百个物种、数千个基因的演化关系?本文以蛙类数据为例,证实了该方案的可行性。我们为138个内类群(ingroup)物种及3784个核标记(超保守元件,ultraconserved elements, UCEs)构建了系统基因组数据集,其中包含70个物种的全新UCE测序数据。此外,我们还组装了一套超级矩阵数据集,覆盖97%的现生蛙属(总计441个属),每个类群包含1至307个基因。随后我们构建了合并后的系统基因组-超级矩阵数据集(即“巨型矩阵(gigamatrix)”),包含441个内类群类群与4091个遗传标记,但整体数据缺失率达86%。对该巨型矩阵进行似然法分析后,得到了总体支持度较高的蛙类类群系统发育树……



