NNS-500 Acceptability judgment task dataset based on the sentences written by non-native English speakers
收藏资源简介:
<strong>Acceptability judgment task (AJT):</strong> AJT is a common method in empirical linguistics to gather information about the internal grammar of speakers of a language, which is considered a promising area to evaluate neural language models' linguistic knowledge. There is a Corpus of Linguistic Acceptability (CoLA) whose creators think Boolean judgments sufficient; similarly, some non-English resources cast acceptability as a binary classification task. <strong>Dataset:</strong> NNS-500 dataset based on the sentences written by non-native speakers (which is important from the point of view of the source of unacceptable sentences) and labelled by a university English teacher is intended for testing the pre-trained neural networks. It has 350 acceptable and 150 unacceptable sentences, which is 70% of acceptability (this compares to 69.2% in the CoLA out-of-domain set). The dataset markup includes standard data: id number, sentence, indication of acceptability – 1, indication of unacceptability – 0, type of error (morphology, syntax, semantics), and detailed information about the source (group number with the year of admission to the university, number according to list of the group members, and gender of the student). For the use of the assessment of EFL learners' linguistic competence, the first 100 sentences of the dataset (id 1‒100) include the ones written by the study group with a high level of academic performance (Group A) and another 100 sentences of the dataset (id 101‒200) are taken from the writing assignments of students with a poor academic performance (Group B). From each group of students, 5 people were selected (a total 10 participants); 20 sentences were randomly selected from each student's written work (14 ‒ unacceptable, 6 ‒ acceptable); there are more sentences with errors, since they are very important for the error analysis. The rest of the dataset consists of 290 acceptable and 10 unacceptable sentences (id 201‒500) taken from the works of students of different study groups of an intermediate level.
可接受性判断任务(Acceptability Judgment Task,AJT):AJT是经验语言学领域中用于采集语言使用者内在语法信息的常用研究范式,同时也是评估神经语言模型语言知识水平的极具潜力的研究方向。语言可接受性语料库(Corpus of Linguistic Acceptability,CoLA)的创建者认为仅需布尔类型的可接受性判断即可满足研究需求;类似地,部分非英语语料资源将可接受性判定任务定义为二分类任务。 数据集:NNS-500数据集基于非母语使用者撰写的语句构建(从不可接受语句的来源维度来看,该构建方式具备独特研究价值),并由高校英语教师完成标注,旨在用于预训练神经网络的性能测试。该数据集共包含350条可接受语句与150条不可接受语句,可接受语句占比达70%,该占比与CoLA域外测试集的69.2%相近。该数据集的标注字段涵盖标准信息项:语句编号(id)、语句文本、可接受性标识(1代表可接受,0代表不可接受)、错误类型(形态学、句法、语义),以及来源详细信息:学生所属年级组与入学年份、组内成员序号、学生性别。 为适配英语作为外语(English as a Foreign Language,EFL)学习者语言能力的评估需求,数据集的前100条语句(编号1至100)均来自学术表现优异的学习小组(A组);另有100条语句(编号101至200)取自学术表现欠佳的学生的写作作业(B组)。研究人员从每个小组中选取5名学生(总计10名参与者),并从每名学生的写作作品中随机抽取20条语句(其中14条为不可接受语句,6条为可接受语句);该部分的不可接受语句占比更高,这是因为错误分析对不可接受语句有较高的需求。数据集剩余部分(编号201至500)包含290条可接受语句与10条不可接受语句,均取自不同年级组中级水平学生的写作作品。



