A Synthetic Data Suite for Evaluating Feature Selection Methods
收藏资源简介:
Subject: Machine Learning / Feature Selection Creators: Athina Bikaki and Ioannis A. Kakadiaris This repository contains a comprehensive synthetic benchmark suite designed for the systematic evaluation and comparison of feature selection algorithms in classification problems. The data are provided in .csv format and include the following four datasets: XOR-25: A binary dataset focused on nonlinearity and interacting features, featuring a small sample-to-feature ratio CLOG-100: A high-dimensional, sparse binary dataset designed to evaluate grouping structures (features form two distinct clusters) CorrAL-100: An expansion of the CorrAL dataset to 100 features, incorporating irrelevant features and noisy copies of relevant ones CIRCLES-190: A real-valued dataset with nonlinear decision boundaries (concentric circles) and 190 features, including trigonometric and PCA-derived distractors



