Toward Controllable Learner Simulation: benchmark data and analysis code
收藏资源简介:
Data and analysis code accompanying the manuscript Toward Controllable Learner Simulation: A Benchmark for Imperfect Student Modeling with Large Language Models (submitted to Education Sciences, MDPI). The benchmark asks whether a large language model can be steered to behave like a student with a specified profile of retained and missing curricular mathematics skills, and measures how faithfully it does so. This deposit contains: Data (imperfectstudent_data.zip): the Grade 4-5 four-option multiple-choice benchmark items with Common Core skill annotations (derived from MathCAMPS), the per-item outputs of Claude, DeepSeek, and GPT-4o across grades and prompt strategies at temperature 0, the prompt templates, and the derived analysis outputs behind the paper's tables and figures. Code (imperfectstudent_code.zip): the analysis pipeline (per-skill controllability metrics, the GEE logistic model, exact McNemar tests, Friedman/Wilcoxon tests, item-level bootstrap, the marginal-preserving permutation null for cross-model error convergence, the cross-skill influence matrices, and the misconception taxonomy) plus figure generation. Benchmark items are derived from the openly released MathCAMPS benchmark and retain its attribution; see EXTERNAL_DATA.md. Model outputs were produced at temperature 0, so the statistics and figures reproduce without any API calls.



