Machine–human agreement in marking undergraduate assignments is specific to model, prompt and institution
收藏资源简介:
This dataset supports the accompanying paper, which evaluates three frontier large language models (Claude, Gemini, GPT) as markers of authentic undergraduate psychology assignments from three UK universities (Cambridge, Manchester Metropolitan, Nottingham). Four de-identified files, linked to a common assignment identifier. assignment_marks.csv - 761 assignments: institution, calibration/test split, human mark, each model's mark under the deployed prompt configuration, four ensemble strategies, word count. calibration_scores.csv - 12,385 factorial marks: 153 assignments × 3 models × 27 cells of a crossed 3 × 3 × 3 prompt design, less eight unreturned Gemini cells. reliability_runs.csv - 2,475 repeated marks (165 assignments, five runs per model). linguistic_features.csv - 33 textual features for the 608 held-out test assignments. While these files allow the full replications of the paper's results, the original assignments are withheld: they are students' personal data and intellectual property, and release would breach the consent granted under Cambridge HE Studies REC approval 2025.LT.84.



