A Comprehensive Collection of 28 Pre-processed Benchmark Datasets for High-Dimensional Small-Sample (HDLSS) Feature Selection
收藏资源简介:
This repository contains 28 classic benchmark datasets specifically curated and pre-processed for research in High-Dimensional Small-Sample (HDLSS) feature selection. Key Features: Unified Format: All datasets are provided in CSV format with labels in the final column. Binary Classification: Multi-class datasets have been converted to binary tasks by selecting the most frequent classes to ensure consistency across benchmarks. Diversity: The collection covers various domains, including microarray (gene expression), mass spectrometry, image recognition, and chemical/biological responses. High-Dimensionality: Includes several extremely high-dimensional datasets such as Arcene (10,000 features), Colon (10,000 features), and GLA_BRA_180 (10,935 features). Source Attribution:The datasets are sourced from the UCI Machine Learning Repository, OpenML, and the ASU Feature Selection Lab. Detailed source information and pre-processing steps are included in the README_DATA.md file within the package.



