遇见数据集

AI-Driven Formative Assessment in EFL Writing: A Comparative Study of ChatGPT-4o and Human Raters

收藏
Zenodo2026-04-20 更新2026-05-26 收录
官方服务:

资源简介:

This study evaluated ChatGPT-4o's potential as a scalable tool for formative assessment in English-as-a-foreign-language (EFL) writing instruction in higher education, with data drawn from a South Korean university context. Using a mixed-methods design that combined Item Response Theory (IRT), qualitative feedback analysis, and the Assessment for Learning (AfL) framework, the study compared ChatGPT-4o's holistic essay scoring and qualitative feedback against those of three experienced, IELTS-certified university English instructors. A total of 76 essays produced by 38 non-English-major undergraduates for the IELTS Academic Writing Test (Task 1 and Task 2) were analyzed during the Fall 2024 semester (September–December 2024). ChatGPT-4o and the human raters scored the essays using the IELTS rubric and provided comments on language use, content quality, and organizational structure. IRT-based results indicate that ChatGPT-4o accounted for 87.63% of person variance compared to 76.59% for human raters, with significantly lower residual variance (10.83% vs. 18.58%), suggesting stronger internal scoring consistency. No statistically significant differences in mean scores were found (F = 0.078, p = 0.972; ICC = 0.792; α = 0.937). Qualitative feedback analysis revealed that ChatGPT-4o delivered more balanced and comprehensive feedback with a strong emphasis on content and organizational structure, whereas teacher feedback focused more heavily on surface-level linguistic accuracy. These findings have implications for AI-assisted writing assessment in large-enrollment EFL contexts globally, where scalable, consistent formative feedback remains a persistent challenge. The datasets supporting the findings of this study are openly available and consist of four files. The primary quantitative file, scores.csv, contains 532 rows recording holistic and sub-criterion scores (Task Achievement, Coherence and Cohesion, Lexical Resource, and Grammatical Range and Accuracy) for each of the 38 student participants across both writing tasks, four raters, and — for ChatGPT-4 — four temporally separated scoring runs. The qualitative dataset, feedback_coding.csv, contains 912 rows representing the full thematic coding of rater feedback, with each row corresponding to one student, one task, one rater, and one of the three feedback categories (Language, Content, Organization). Binary flags for surface-level focus, actionability, and higher-order concern are included for each coded instance. The psychometric output file, irt_parameters.csv, provides the Many-Facet Rasch Model estimates for all four raters, including severity, infit and outfit mean-square statistics, person variance decomposition, and aggregate reliability indices as reported in Tables 1, 3, and 4 of this paper. All variable definitions, value ranges, coding procedures, theme code descriptions, and inter-coder reliability notes are documented in the accompanying codebook.pdf. Student identifiers are fully anonymized. Data were collected under written informed consent.

本研究以韩国某高校为研究场景,探讨了ChatGPT-4o作为可规模化工具应用于高等教育英语作为外语(English-as-a-foreign-language, EFL)写作形成性评估的潜力。研究采用混合研究设计,整合了项目反应理论(Item Response Theory, IRT)、质性反馈分析与学习型评估(Assessment for Learning, AfL)框架,将ChatGPT-4o的整体作文评分与质性反馈与三位具备雅思(IELTS)认证的资深大学英语教师的评分及反馈进行对比。2024年秋季学期(2024年9月—12月)期间,研究者共分析了38名非英语专业本科生完成的76篇雅思学术类写作考试(任务1与任务2)作文。ChatGPT-4o与人类评分员均采用雅思评分细则对作文进行评分,并针对语言使用、内容质量与组织结构提供评论。基于项目反应理论的结果显示,ChatGPT-4o可解释87.63%的个体方差,而人类评分员的解释度为76.59%;前者的残差方差显著更低(10.83% vs. 18.58%),表明其评分内部一致性更强。两组平均评分无统计学显著差异(F = 0.078, p = 0.972; 组内相关系数(Intraclass Correlation Coefficient, ICC)= 0.792; 克朗巴赫α系数(Cronbach's α)= 0.937)。质性反馈分析结果表明,ChatGPT-4o提供的反馈更为均衡全面,且重点突出内容与组织结构;而教师反馈则更侧重表层语言准确性。上述研究发现对全球范围内大班额EFL语境下的AI辅助写作评估具有启示意义——此类场景中,可规模化、一致性强的形成性反馈始终是一项亟待解决的难题。 本研究的支撑性数据集已公开,共包含4个文件。主量化数据文件scores.csv包含532条记录,涵盖38名学生在两项写作任务中的整体评分与分项评分(任务完成度、连贯与衔接、词汇资源、语法范围与准确性),涉及4位评分员,且针对ChatGPT-4的评分还包含4次不同时间的评分结果。质性数据集feedback_coding.csv包含912条记录,对应所有评分员反馈的完整主题编码,每条记录对应一名学生、一项写作任务、一位评分员,以及三类反馈类别(语言、内容、组织结构)中的一类。每条编码实例均包含表层关注点、可操作性与高阶关注点的二进制标记。心理测量学输出文件irt_parameters.csv提供了四位评分员的多面拉什模型(Many-Facet Rasch Model)估计值,包括评分严格性、infit与outfit均方统计量、个体方差分解以及汇总信度指标,与本文表1、表3和表4中报告的内容一致。 所有变量定义、取值范围、编码流程、主题代码说明以及编码者间信度说明均在配套的codebook.pdf文件中进行了记录。学生身份信息已完全匿名化。数据收集过程已获得书面知情同意。

提供机构:
Zenodo
创建时间:
2026-04-19
二维码
社区交流群
二维码
科研交流群
商业服务