遇见数据集

EduGen-Dataset

收藏
Zenodo2026-02-22 更新2026-05-26 收录
官方服务:

资源简介:

EduGen Curated Student Performance Dataset (v1.1) The EduGen Curated Student Performance Dataset (v1.1) is a harmonized, feature-engineered educational dataset constructed to support reproducible research in adaptive assessment, learner profiling, psychometric modeling, and AI-driven educational analytics. The dataset was curated by integrating and preprocessing two publicly available Kaggle datasets: Students Grading Dataset Student Performance and Clustering Dataset The original datasets provide learner-level academic records, including quiz scores, attendance metrics, behavioral indicators, and performance attributes. Historical Computational Training Dataset To ensure consistency and suitability for adaptive modeling, the datasets were schema-aligned, normalized, and enriched with engineered features, including: Consistency_Index (standard deviation of quiz scores across Q1–Q12) Top_9_Sum (aggregation of best nine quiz scores) Normalized_Final_Score Performance_Class ∈ {High, Medium, Low} Behavioral and engagement indicators No original labels were modified. All additional variables were derived solely for modeling, benchmarking, and evaluation purposes. This historical dataset contains 1,053 learner records and supports supervised classification, clustering, psychometric calibration (IRT), and adaptive difficulty modeling. Empirical Pilot Dataset (Adaptive vs Static Assessment Study) Version 1.1 additionally includes an anonymized Empirical Pilot Dataset (N = 80) collected from a controlled classroom study evaluating adaptive assessment effectiveness. Participants were randomly assigned to: Control Group (Static Quiz) – N = 40 Experimental Group (EduGen Adaptive Workflow) – N = 40 The pilot dataset contains the following attributes: Participant_ID Group_ID (Control / Experimental) Quiz_Score_Norm (0–100%) Q_Difficulty_Var (variance of encountered question difficulty) Time_On_Task (minutes) Adaptivity_Score (ordinal 1–5 perception scale) Hallucination_Flag (0/1; instructor-verified factual inconsistency) This dataset enables empirical validation of: Adaptive difficulty sequencing Pedagogical alignment Hallucination reduction through retrieval grounding Learning efficiency (time-aware assessment) Perceived adaptivity metrics All pilot data are anonymized and contain no personally identifiable information. Intended Research Applications The combined dataset (Historical + Pilot) supports research in: Adaptive learning systems Learner state estimation Item Response Theory (2PL modeling) Bloom-level cognitive alignment Retrieval-Augmented Generation (RAG) in education Benchmark and scholarship suitability modeling Educational AI evaluation metrics (PSR, calibration, hallucination rate) Reproducibility The repository includes: Processed datasets Feature engineering documentation Variable specifications Preprocessing scripts Train/validation/test split configuration (70/15/15) The dataset is released to promote transparent, reproducible, and benchmarkable research in AI-powered adaptive education systems.

提供机构:
Zenodo
创建时间:
2026-02-22
二维码
社区交流群
二维码
科研交流群
商业服务