Simulated Dataset: wQFM-GDL Enables Accurate Quartet-based Genome-scale Species Tree Inference Under Gene Duplication and Loss
收藏资源简介:
We simulated two large datasets, SIM200 and SIM500 containing 200 and 500 taxa model conditions, respectively to test our developed method wQFM-GDL against other popular methods. For both datasets, we varied three duplication rates. For each duplication rate, we considered two loss rates, two levels of ILS, and different numbers of gene trees (250, 500, and 1000). This resulted in a total of 36 model conditions per dataset, with 20 replicates for SIM200 and 10 replicates for SIM500. For both SIM200 and SIM500, We used Simphy to generate true gene family trees and AliSim to simulate sequence alignment from the true trees. Finally, FastTree was used to estimate maximum likelihood gene trees using the GTR+Gamma model. We varied duplication and loss rates, ILS level, and the number of gene trees for both of the datasets. The details of the parameters are presented in Supplementary Table 1. The proper duplication rates and haploid effective population sizes for GDL and ILS levels of the model conditions are estimated empirically. The dataset includes the estimated multicopy gene family trees in the file estimated_gene_trees.zip, and the true gene trees simulated by SimPhy along with their corresponding alignments in true_trees_and_MSA.zip. Within each zip archive, the data are organized into subfolders by model condition, and each of these contains additional subfolders for individual replicates.



