Performance of algorithms (area under ROC curve, AUC) in the rediscovery experiment using only NEM316 genome.
收藏资源简介:
This analysis evaluated the relative performance of each algorithm to rediscover virulence genes by applying stratified n-fold cross-validations with of the entire set of S. agalactiae NEM316 genes serving as test-set in each fold. Each fold of training set comprised positive and negative examples.n: number of virulence genes in the category. Singleton virulence gene categories were excluded from this analysis, as it is not possible to perform cross-validations on training sets with n = 1. All but one (labeled*) AUCs reached the statistical significance level at α = 0.05 (two-tailed Mann-Whitley U-test). At least 3 out of 4 algorithms were still significant after adjustment for multiple testing (across the family of 4 algorithms) by the Bonferroni method. Abbreviations: ADTree: alternating decision tree; IBk: nearest neighbor classifier; SVM: support vector machine; RBF: SVM with radial basis function; Poly: SVM with polynomial kernel. Refer to the methods section for the parameters used to train the machine learning algorithms. The numbers in bold face indicate the best performing algorithm for a given category.
本分析评估了各算法通过分层n折交叉验证重新发掘毒力基因的相对性能,其中每一折均以无乳链球菌(S. agalactiae)NEM316的全部基因集合作为测试集。每一折的训练集均包含正样本与负样本。n:该类别中毒力基因的数量。仅含单个毒力基因的类别被排除在本次分析之外,因无法对n=1的训练集执行交叉验证。除标记为*的1项指标外,所有受试者工作特征曲线下面积(Area Under Curve, AUC)均达到α=0.05的统计学显著性水平(双侧曼-惠特尼U检验)。经邦费罗尼(Bonferroni)法对4种算法家族进行多重检验校正后,仍至少有3种算法保持统计学显著性。缩写说明:ADTree:交替决策树(alternating decision tree);IBk:近邻分类器(nearest neighbor classifier);SVM:支持向量机(support vector machine);RBF:带径向基核函数的支持向量机(SVM with radial basis function);Poly:带多项式核函数的支持向量机(SVM with polynomial kernel)。机器学习算法的训练参数详见方法部分。粗体数字代表对应类别中性能最优的算法。



