Empirical Dataset for Evaluating a Generative-AI Essay Grading Plugin in Moodle
收藏资源简介:
The data consist of: (1) 250 student essay responses spanning five subjects (Indonesian Language, History, Geography, Biology, Mathematics), each scored by experienced teachers and by the DeepSeek-R1 model to compute agreement metrics (e.g., Pearson r, MAE); (2) performance/throughput logs capturing end-to-end grading time across different answer lengths and simulated class sizes to estimate efficiency gains versus manual grading; (3) a four-week field deployment in a Bali school with a survey of 25 teachers using 5-point Likert items to assess acceptability, ease of use, and perceived utility; and (4) ablation data from prompt-engineering experiments (basic, chain-of-thought, rubric-based, few-shot, and comprehensive prompts) to compare accuracy across designs and question types.



