ASP-POTASSCO
收藏资源简介:
This dataset was curated for [TabArena](https://tabarena.ai/) by the TabArena team as part of the [TabArena Tabular ML IID Study](https://tabarena.ai/data-tabular-ml-iid-study). For more details on the study, see our [paper](https://tabarena.ai/paper-tabular-ml-iid-study). **Dataset Focus**: This dataset shall be used for evaluating predictive machine learning models for independent and identically distributed tabular data. The intended task is classification. --- #### Dataset Metadata - **Licence:** Public - **Original Data Source:** https://www.openml.org/search?type=data&sort=runs&status=active&id=41705 - **Reference (please cite)**: Hoos, Holger, Marius Lindauer, and Torsten Schaub. 'claspfolio 2: Advances in algorithm selection for answer set programming.' Theory and Practice of Logic Programming 14.4-5 (2014): 569-585. https://doi.org/10.1017/S1471068414000210 - **Dataset Year:** 2014 - **Dataset Description:** see the reference and the original data source for details. #### Curation comments by the TabArena team (for code see the [page of the study](https://tabarena.ai/data-tabular-ml-iid-study)): - We dropped the ID column. - We dropped the constant column "repetition". - We drop both "runtime" and "runstatus" as they are leaking the target. Both can only be computed after the algorithm has been run. But the task is to predict which algorithm to run. We also drop all instances where "timeout" is not "ok" as this would be the false target to predict.




