遇见数据集

A Multilingual Dataset of Student Answers, Human Grading, and Multi-LLM Evaluations for Automated Assessment Research using JorGPT

收藏
Zenodo2026-03-24 更新2026-05-26 收录
官方服务:

资源简介:

This is a multilingual dataset of open-ended student answers collected from authentic university-level assessments. The dataset comprises 3041 student responses to 50 distinct open-ended questions, authored by 79 anonymized students from a single institution across one semester of the academic course 2025-26. It integrates real student answers, instructor-defined ideal solutions, numerical grades and qualitative feedback provided by human teachers, and structured evaluations generated by multiple state-of-the-art LLMs following a unified grading schema. All pedagogical content is available in both Spanish, the original language of instruction, and English, enabling multilingual and cross-lingual research. The dataset supports benchmarking of automated grading systems, analysis of alignment between human and LLM-based assessment, training of judge or meta-evaluation models, and studies on automated feedback generation. By releasing this dataset under an open license, we aim to facilitate transparent, reproducible, and realistic research in educational artificial intelligence and automated assessment. The full description of the characteristics of the dataset, as well as its components, format and methodology, can be found here: "Multilingual Dataset of Student Answers, Human Grading, and Multi-LLM Evaluations for Automated Assessment Research Using JorGPT". Data 2026, 11, 59. https://doi.org/10.3390/data11030059 The dataset is also available on Kaggle: https://www.kaggle.com/datasets/javiersanchezsoriano/jorgpt-student-answers-and-multi-llm-grading

提供机构:
Zenodo
创建时间:
2026-01-25
二维码
社区交流群
二维码
科研交流群
商业服务