Additional file 9. Sheet 1. The data used for running random forest machine learning model. Only the coding sequences (CDS) that were deleted in all isolates of a phylogenetic cluster are used for thi
This data accompanies the manuscript "Cross-platform normalization enables machine learning model training on microarray and RNA-seq data simultaneously" by Foltz, Taroni, and Greene. Please refer to