TabReD
收藏资源简介:
TabReD是由Yandex和HSE大学创建的一个包含八个工业级表格数据集的基准,覆盖金融、食品配送等多个领域。这些数据集具有时间分割特性,支持基于时间序列的训练和测试分割,反映了真实世界数据的时间演变特性。数据集通过从Kaggle竞赛和工业ML应用中收集,经过精心筛选和特征工程处理,确保了数据的质量和实用性。TabReD主要用于评估和推动表格机器学习模型的发展,特别是在处理时间序列数据和复杂特征工程方面的应用。
TabReD is a benchmark comprising eight industrial-grade tabular datasets developed by Yandex and HSE University, covering a wide range of domains such as finance and food delivery. These datasets possess time-splitting characteristics, enabling time-series-based training and test partitioning, which mirror the temporal evolution patterns inherent in real-world data. Collected from Kaggle competitions and industrial machine learning applications, the datasets have been meticulously screened and processed with feature engineering to guarantee their data quality and practical applicability. TabReD is primarily designed to evaluate and advance the progress of tabular machine learning models, particularly for scenarios involving time-series data and complex feature engineering tasks.

- 1TabReD: A Benchmark of Tabular Machine Learning in-the-WildYandex, HSE大学 · 2024年



