Exhaustive k-nearest-neighbour error tables over all feature subsets of 12 UCI-derived benchmark datasets
收藏资源简介:
Misclassification counts for every non-empty feature subset of 12 UCI-derivedbenchmark datasets with at most 22 features, computed by exhaustive enumerationwith a k-nearest-neighbour classifier (k = 5, Euclidean distance, min-maxscaling fitted on the training part only, fixed tie rules). The deposit holds 120 tables: 12 datasets x 5 stratified outer folds x 2 splits.For each fold there is a validation table (inner-train to inner validation) anda test table (outer-train to outer test). Each table stores one unsigned 16-biterror count per subset, indexed by the subset bitmask, from 511 subsets for the9-feature datasets to 4,194,303 for the 22-feature SpectEW. The empty subset isstored as 65535 and is excluded from every analysis. The tables were checked against an independent plain-numpy implementation of thesame rules on 199,703 subsets with no mismatch, and building them took 984seconds in total on an Apple M2 MacBook Air. These tables give the exact optimum of a stated objective on a stated split, soany feature selection method can be scored against a checked number rather thanagainst the best value a competitor happened to find. The code that builds,verifies and uses them, together with the benchmark results, is depositedseparately. Licence: CC BY 4.0. The underlying data files derive from the UCI MachineLearning Repository and are credited in DATA-ATTRIBUTION.md in the accompanyingsoftware deposit.



