遇见数据集

The HARMOGEN-R dataset of human and AI-generated rubric evaluations for formative programming assessment

收藏
Zenodo2025-10-08 更新2026-05-26 收录
官方服务:

资源简介:

HARMOGEN-R Dataset: Large-Scale Human–LLM Evaluation Records for Automated Grading Research The HARMOGEN-R dataset comprises 69,028 structured records of educational assessments that compare human and Large Language Model (LLM) grading. It includes 770 anonymized student responses from a Data Structures and Algorithms course, evaluated by human instructors and several LLMs using five rubric variants: one human-created and four AI-generated. Each response was assessed across multiple criteria, producing 50,050 individual evaluations, 15,480 aggregated scores, and 2,310 evaluation reports. The dataset consists of 15 relational tables with enforced foreign keys and validated JSON fields that document the reasoning process used in automated grading. Each LLM evaluation includes a criteria_evaluations_by_llm JSON object describing the structured rationale applied by the model to each criterion, the intermediate sub-scores, and the corresponding textual justification. This structure enables reproducible quantitative and qualitative analyses of model evaluation behaviour. The repository also provides CSV extracts, Python scripts for data normalization and exploratory analysis, and MySQL schema files to ensure reproducibility. All identifiers are anonymized, and the dataset does not contain personal information. The HARMOGEN-R dataset facilitates empirical research on automated assessment, rubric generation, and human–AI grading consistency. It is released under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.

提供机构:
Zenodo
创建时间:
2025-10-08
二维码
社区交流群
二维码
科研交流群
商业服务