遇见数据集

An Expert-Annotated IELTS Writing Corpus with Trait-Level Scores and Fine-Grained Feedback for Educational NLP

收藏
Zenodo2025-07-05 更新2026-05-26 收录
官方服务:

资源简介:

Writing assessment is foundational to education, language proficiency evaluation, and increasingly, to the benchmarking oflarge language models (LLMs) on open-ended text generation tasks. Despite growing interest, the development of robust,explainable evaluation systems has been limited by the lack of large-scale, high-quality, and rubric-aligned writing datasets. Toaddress this gap, we introduce a comprehensive IELTS writing corpus designed to support both automated essay scoring(AES) and educational feedback generation. The dataset comprises two complementary subsets. The first includes 4,200expert-annotated essays—curated from an initial pool of 5,088 responses—with trait-level scores based on the official IELTSrubric (Task Response, Coherence and Cohesion, Lexical Resource, Grammatical Range and Accuracy) and detailed, span-level corrective feedback. The second subset contains 80,800 unannotated essays sourced from learner communities andpreparation platforms, offering a rich resource for pretraining, instruction tuning, and model evaluation. Annotation proceduresinvolved certified IELTS instructors, double-blind scoring, inter-rater reliability (Fleiss’ Kappa ≥ 0.80), and GPT-4-assistedfeedback validation. This dataset provides a valuable foundation for building interpretable, criterion-aligned writing assessmentsystems and for advancing the evaluation of LLMs in educational and NLP research contexts.

提供机构:
Zenodo
创建时间:
2025-07-05
二维码
社区交流群
二维码
科研交流群
商业服务