遇见数据集

Toward Controllable Learner Simulation: benchmark data and analysis code

收藏
Zenodo2026-08-08 更新2026-08-13 收录
官方服务:

资源简介:

Data and analysis code accompanying the manuscript Toward Controllable Learner Simulation: A Benchmark for Imperfect Student Modeling with Large Language Models (submitted to Education Sciences, MDPI). The benchmark asks whether a large language model can be steered to behave like a student with a specified profile of retained and missing curricular mathematics skills, and measures how faithfully it does so. This deposit contains: Data (imperfectstudent_data.zip): the Grade 4-5 four-option multiple-choice benchmark items with Common Core skill annotations (derived from MathCAMPS), the per-item outputs of Claude, DeepSeek, and GPT-4o across grades and prompt strategies at temperature 0, the prompt templates, and the derived analysis outputs behind the paper's tables and figures. Code (imperfectstudent_code.zip): the analysis pipeline (per-skill controllability metrics, the GEE logistic model, exact McNemar tests, Friedman/Wilcoxon tests, item-level bootstrap, the marginal-preserving permutation null for cross-model error convergence, the cross-skill influence matrices, and the misconception taxonomy) plus figure generation. Benchmark items are derived from the openly released MathCAMPS benchmark and retain its attribution; see EXTERNAL_DATA.md. Model outputs were produced at temperature 0, so the statistics and figures reproduce without any API calls.

提供机构:
Zenodo
创建时间:
2026-08-08
二维码
社区交流群
二维码
科研交流群
商业服务