ASAP dataset
收藏资源简介:
ASAP数据集是由Kaggle平台主办的自动学生评估竞赛所构建的英文作文评分基准数据集,由学术机构联合教育部门采集编纂。该数据集共包含12,978篇学生作文,涵盖议论文、源文本依赖型和叙事文三种写作类型,涉及八个不同的写作提示,每篇作文均附有人工评定的分数标签,作文平均长度在150至650词之间。数据集的构建过程通过标准化教育评估流程收集真实学生作文,并由专业评分员根据统一量规进行多维度评分。该数据集主要应用于教育技术领域的自动作文评分和反馈生成研究,旨在通过机器学习方法实现对学生写作能力的自动化评估,减轻教师批改负担并提供个性化学习指导。
The ASAP dataset is a benchmark dataset for English essay scoring developed for the Automated Student Assessment Prize (ASAP) competition hosted on the Kaggle platform, and it was collected and curated by a consortium of academic institutions and educational authorities. This dataset comprises 12,978 student essays in total, covering three writing categories: argumentative essays, source-text-dependent essays, and narrative essays, associated with eight distinct writing prompts. Each essay is equipped with manually annotated score labels, and the average length of the essays ranges from 150 to 650 words. The dataset is constructed by gathering real student essays via standardized educational assessment workflows, with professional raters performing multi-dimensional scoring based on unified rubrics. This dataset is primarily applied to research on automated essay scoring and feedback generation in the field of educational technology, with the goal of realizing automated assessment of students' writing abilities through machine learning methods, thus alleviating teachers' grading burden and providing personalized learning guidance.
数据集概述
- 项目名称:Beyond Score Prediction: LLM-Based Essay Scoring and Feedback Generation via Reinforcement Learning with Rubric Rewards
- 论文地址:https://arxiv.org/abs/2607.19219
- 当前状态:代码与数据集将在论文被接收后发布。
- 引用格式:详见项目页面中的 BibTeX 引用信息。

- 1Beyond Score Prediction: LLM-Based Essay Scoring and Feedback Generation via Reinforcement Learning with Rubric Rewards阿里巴巴集团·阿里云 · 2026年



