遇见数据集

Question-Level Human–AI Grading Dataset for Undergraduate Computing Courses (IT2430/IT2431/IT1230, Spring 2024–Fall 2024)

收藏
Zenodo2026-02-26 更新2026-05-26 收录
官方服务:

资源简介:

This dataset contains paired, question-level grading records used to evaluate agreement, bias, and calibration between instructor (human) grading and AI-assisted grading in undergraduate computing education. It includes 2,480 graded question-response records collected from three Georgia Southern University courses: IT2430 (Data Programming I), IT2431 (Computer Programming II), and IT1230 (Web Technologies), spanning two academic terms and 120 unique questions across 24 instructional topics. Each row represents one student response to one question and provides both the human-assigned and AI-generated scores on the same item, along with contextual metadata. Scores are reported on the original point scale of the item (max_points = 5 or 10), enabling normalization and direct paired comparisons. Columnscourse_id: Course identifier (e.g., IT2430)term: Academic term (e.g., Spring 2024)question_id: Unique question identifiertopic: Instructional topic labelmax_points: Maximum possible points for the question (5 or 10)human_score: Human-assigned scoreai_score: AI-generated scorehuman_feedback_length_chars: Length of human feedback in charactersai_feedback_length_chars: Length of AI feedback in characters Intended useThe dataset supports reproducible analyses of human–AI grading alignment at question-level resolution (e.g., correlation, ICC, weighted kappa over binned grades, Bland–Altman diagnostics, MAE/RMSE, and decile-based calibration). It is referenced in the companion paper “Human–AI Grading Agreement in Computing Education: A Question-Level Reliability and Calibration Study.”

提供机构:
Zenodo
创建时间:
2026-02-26
二维码
社区交流群
二维码
科研交流群
商业服务