TILBench
收藏资源简介:
TILBench是由苏州大学研究团队构建的一个大规模表格不平衡学习基准数据集,旨在系统评估不同算法在多样化数据特征下的性能。该数据集汇集了57个表格分类任务,涵盖二元与多元分类,包含不同规模、特征维度、不平衡比率及缺失值的数据,数据主要来源于OpenML和imbalanced-learn开源平台。其创建过程通过统一且可复现的评估框架,整合了超过40种代表性算法,进行了超过20万次受控实验。该数据集的应用领域聚焦于解决现实世界中的表格数据不平衡学习问题,如欺诈检测、医疗诊断等,为在不同数据特性与计算约束下选择合适方法提供实证依据与实用指南。
TILBench is a large-scale tabular imbalanced learning benchmark dataset constructed by the research team from Soochow University, which aims to systematically evaluate the performance of diverse algorithms under varied data characteristics. This dataset encompasses 57 tabular classification tasks covering binary and multi-class classification, including data with different scales, feature dimensions, imbalance ratios, and missing values. The data is primarily sourced from the open-source platforms OpenML and imbalanced-learn. During its development, a unified and reproducible evaluation framework was employed, integrating over 40 representative algorithms and conducting more than 200,000 controlled experiments. The application domains of this dataset focus on addressing real-world tabular imbalanced learning problems such as fraud detection and medical diagnosis, providing empirical evidence and practical guidelines for selecting appropriate methods under different data characteristics and computational constraints.

- 1TILBench: A Systematic Benchmark for Tabular Imbalanced Learning Across Data Regimes苏州大学·数学科学学院 · 2026年



