遇见数据集

Unfair Inequality in Education: A Benchmark for AI-Fairness Research (Aequitas WP7 Use Case S2)

收藏
Zenodo2025-05-29 更新2026-05-26 收录
官方服务:

资源简介:

Unfair Inequality in Education: A Benchmark for AI-Fairness Research This dataset proposes a novel benchmark specifically designed for AI fairness research in education. It can be used for challenging tasks aimed at improving students' performance and reducing dropout rates which are also discussed in the paper to emphasize significant research directions. By prioritizing fairness, this benchmark aims to foster the development of bias-free AI solutions, promoting equal educational access and outcomes for all students. Structure benchmark contains: the proposed dataset (dataset.csv), the mask for dealing with missing values (missing_mask.csv), and the meta-columns providing grouping criteria and sample weights for each student (meta_cols.csv). raw_data includes: the original dataset (original.csv), and the intermediate stages of the pre-processing and validation pipelines (split, pre_processed, and validation). res contains the documentation, including: the transformation mapping each column of the original dataset to the proposed one, along with the missingness category and original text (meta_data_mapping.csv), the value type and domains of each column of the proposed datasets (meta_data_stats.json), and the statistical indices of the validation pipeline (bias_preservation_results.json). src contains the source code for running the pre-processing and corresponding analysis: pre_processing and statscontain the code for the two corresponding tasks, and pre_processing.py and split.py are two entry points. Finally, Dockerfile and requirements.txt set up the environment for running the applications across multiple platforms and with Python, respectively.

教育领域的不公平差距:面向AI公平性(AI Fairness)研究的基准数据集 本数据集专为教育场景下的AI公平性研究打造,是一款全新的基准测试集。其可用于攻克提升学生学业表现、降低辍学率等挑战性任务——本文亦对这类任务展开探讨,以明确核心研究方向。本基准测试集以公平性为首要导向,旨在推动无偏AI解决方案的研发,助力实现全体学生受教育机会与学业成果的均等化。 结构 基准测试集包含以下内容: 本次构建的数据集(dataset.csv), 用于处理缺失值的掩码文件(missing_mask.csv),以及 为每位学生提供分组标准与样本权重的元列文件(meta_cols.csv)。 原始数据目录(raw_data)包含: 原始数据集(original.csv),以及 预处理与验证流水线的中间产物(split、pre_processed及validation)。 结果与文档目录(res)包含以下文档材料: 原始数据集各列到本次构建数据集的转换映射表,附缺失值类别与原始文本说明(meta_data_mapping.csv), 本次构建数据集各列的数据类型与取值范围说明(meta_data_stats.json),以及 验证流水线的统计指标结果(bias_preservation_results.json)。 源代码目录(src)包含用于运行预处理与对应分析的源代码: pre_processing与stats目录分别对应两项任务的代码,且 pre_processing.py与split.py为两个程序入口文件。 最后,Dockerfile与requirements.txt分别用于配置跨平台运行环境与Python依赖环境。

提供机构:
Zenodo
创建时间:
2025-05-29
二维码
社区交流群
二维码
科研交流群
商业服务