Supplementary Data: Mpboot: Fast Phylogenetic Maximum Parsimony Tree Inference And Bootstrap Approximation
收藏资源简介:
Supplementary Data<br> MPBoot: Fast phylogenetic maximum parsimony tree inference and bootstrap approximation<br> Submitted to BMC Evolutionary Biology This record contains PANDIT based dataset and TreeBASE dataset (Nguyen et al. 2015) which are analyzed by different bootstrap methods in the study "MPBoot: Fast phylogenetic maximum parsimony tree inference and bootstrap approximation". The PANDIT based dataset (compressed in file data_pandit.tar.gz) is used to benchmark the accuracy of bootstrap estimates. The TreeBASE dataset (compressed in file data_treebase.tar.gz) is used to benchmark computing times and capability of finding the best-known MP scores. After being uncompressed, the PANDIT based dataset comprises two subdirectories corresponding to the simulated DNA and AA MSAs. They were generated by Seq-Gen (Rambaut and Grass 1997), where the model parameters and true tree were inferred from the original MSAs downloaded from the PANDIT database (Whelan et al. 2006). Inside "dna" subdirectory, there are 6,207 numbered directories corresponding to 6,207 DNA MSAs. Note that the numbering of these directories is not consecutive because we excluded MSAs where TNT or PAUP* runs did not finish. In each numbered directory N, there are three files: (1) data.N contains the simulated MSA in PHYLIP format; (2) model.N contains the best-fit model detected from the corresponding original MSA; (3) tree.N contains the tree (in Newick format) inferred from the corresponding original MSA. tree.N and model.N are used by Seq-Gen to simulate the MSA in data.N. The "aa" subdirectory is organized similarly for 6,165 AA MSAs. After being uncompressed, the TreeBASE dataset comprises 115 files corresponding to 115 MSAs. There are: 70 DNA MSAs in PHYLIP format. These files follow the naming scheme dna_[number of sequences]_[number of sites].phy. 45 protein MSAs in PHYLIP format. These files follow the naming scheme prot_[number of sequences]_[number of sites].phy.<br>



