Student Performance Prediction Dataset.
收藏资源简介:
This dataset contains the research-ready, preprocessed version of the publicly available Student Performance dataset originally compiled by Cortez and Silva (2008). The original dataset includes academic achievement records from two Portuguese secondary schools, along with demographic, socio-economic, behavioral, and school-related attributes. The raw data used in this research was obtained from Kaggle at: https://www.kaggle.com/datasets/henryshan/student-performance-prediction . For the purposes of the study, several preparation steps were applied to produce the research-ready version uploaded here. These steps include data cleaning, label preparation, feature selection, transformation of academic grades (G1, G2, G3), harmonization of categorical variables, and formatting into a machine-learning friendly structure. No synthetic information or additional records were introduced. The uploaded dataset reflects exactly the data used in the final experiments of the study, enabling full reproducibility of the results. This dataset supports the article “Explainable Machine Learning for Student Academic Performance Prediction in Data-Constrained Educational Settings.” The authors of the study do not claim ownership of the original dataset. All rights, authorship, and credit for the original data belong to the original creators and the Kaggle uploader. This upload serves solely to preserve the exact version of the dataset used for reproducibility in accordance with open-science practices. keywords: Student Performance Prediction, Machine Learning, Fairness Auditing, Interpretability Stability, Dataset Constraint. Original Source Credit: Cortez, P., & Silva, A. (2008). “Using Data Mining to Predict Secondary School Student Performance.” Available through Kaggle at: https://www.kaggle.com/datasets/henryshan/student-performance-prediction
本数据集为公开可用的学生成绩数据集(Student Performance Dataset)经预处理后的可供研究使用版本,该原始数据集由Cortez与Silva于2008年汇编完成。原始数据集涵盖两所葡萄牙中学的学业成绩记录,同时包含人口统计学、社会经济、行为学及学校相关属性信息。本研究所用原始数据获取自Kaggle平台,链接为:https://www.kaggle.com/datasets/henryshan/student-performance-prediction。 为适配本研究需求,我们对原始数据实施了多步预处理流程,以生成本次上传的可供研究使用的数据集版本。具体预处理步骤包括数据清洗、标签构建、特征选择、学业成绩(G1、G2、G3)转换、分类变量统一化处理,以及格式化为机器学习友好的数据结构。本次预处理未引入任何合成数据或额外记录,上传的数据集完整复刻了本研究最终实验所用的全部数据,可支撑研究结果的完全复现。 本数据集用于支撑论文《面向数据受限教育场景的学生学业成绩预测可解释机器学习》。本研究作者并不主张对原始数据集拥有所有权,原始数据的全部版权、署名权及相关赞誉均归属于原始创作者及该Kaggle数据集的上传者。本次上传仅为遵循开放科学实践,留存本研究用于复现的精准数据集版本。 关键词:学生成绩预测、机器学习、公平性审计、可解释稳定性、数据集约束 原始来源致谢: Cortez, P. 与 Silva, A. (2008). 《利用数据挖掘预测中学学生学业表现》。该数据集可通过Kaggle平台获取,链接为:https://www.kaggle.com/datasets/henryshan/student-performance-prediction




