A Suite of Fairness Datasets for Tabular Classification
收藏资源简介:
本数据集由IBM研究院和卡内基梅隆大学合作创建,包含20个用于评估机器学习分类器公平性的表格数据集。数据集大小和数据量各异,主要来源于OpenML、AHRQ和ProPublica等平台。创建过程中,数据集经过最小化预处理,并提供公平性元数据,如有利标签和受保护属性。这些数据集主要应用于机器学习公平性研究,旨在通过严格的实验评估,帮助研究者和利益相关者选择和发明更公平的算法。
This dataset was collaboratively developed by IBM Research and Carnegie Mellon University, consisting of 20 tabular datasets for evaluating the fairness of machine learning classifiers. The datasets vary in scale and data volume, and are primarily sourced from platforms such as OpenML, AHRQ, and ProPublica. During the creation process, the datasets underwent minimal preprocessing, with fairness-related metadata including favorable labels and protected attributes provided. These datasets are mainly utilized in machine learning fairness research, aiming to help researchers and stakeholders select and devise fairer algorithms through rigorous experimental evaluations.




