Datasets for "Models trained with noisy genomes extend bacterial phenotype prediction into deep time"
收藏资源简介:
Datasets for training models on annotated genomes for five phenotypes: oxygen use, cell envelope, sporulation, optimal growth temperature, and GC content. data_[phenotype].tar.gz contains the following sub-directories input_data/ contains sub-directories for each taxonomy level and the taxonomy-aware train/test splits, in 80/20 proportion (30 splits for each tax level), input_data_train_val_test/ contains sub-directories for each taxonomy level and the taxonomy-aware train/validation/test splits, in 60/20/20 proportion (30 splits for each tax level), outputs/ contains outputs for each taxonomy level. data_preparation.tar.gz contains the GTDB metafiles with genome taxonomies that are used to create taxonomy-aware train/test splits. phylo_trees.tar.gz contains the GTDB tree files used in the phylogeny-based phenotype prediction algorithm.



