遇见数据集

A harmonized longitudinal dataset of engineering students for Learning Analytics and Predictive modeling

收藏
Mendeley Data2026-05-21 收录
官方服务:

资源简介:

This dataset contains 1,656 student records with 56 variables, synthesized from academic administrative data at an application-oriented university. The data were compiled from multiple spreadsheet-based sources across several academic years and underwent cleaning, standardization, and harmonization to ensure consistency in naming conventions and coding structures. The dataset preserves real-world curriculum variations, including changes in course naming, sequencing, and elective pathways. It is provided in two formats: a labeled version for interpretation and a harmonized coded version for analysis. Variables include socio-demographic information, course grades, cumulative academic performance indicators, scholarship records, and academic status. The dataset can be used for learning analytics, educational data mining, machine learning applications (e.g., student performance prediction), and research on data integration and synthetic data validation.

本数据集包含1656条学生记录,共涉及56个变量,其数据源自某应用型大学的学术行政数据并经合成生成。该数据集的原始数据取自覆盖多个学术年度的多份电子表格类数据源,历经数据清洗、标准化与协调统合处理,以确保命名规范与编码结构的一致性。 本数据集保留了真实的课程设置差异,包括课程命名、修读顺序与选修路径的变动。该数据集提供两种格式:其一为带标注版本,用于数据解读;其二为统合编码版本,用于数据分析。 数据集的变量涵盖社会人口统计学信息、课程成绩、累计学业表现指标、奖学金记录以及学籍状态。本数据集可用于学习分析、教育数据挖掘、机器学习应用(如学生成绩预测)以及数据集成与合成数据验证相关研究。

创建时间:
2026-06-01
二维码
社区交流群
二维码
科研交流群
商业服务