iRead4Skills Dataset 2: annotated corpora by level of complexity for FR, PT and SP
收藏资源简介:
The iRead4Skills Dataset 2: annotated corpora by level of complexity for FR, PT and SP is a collection of texts categorized by complexity level and annotated for complexity features, presented in xlsx format. These corpora were compiled, classified and annotated under the scope of the project iRead4Skills – Intelligent Reading Improvement System for Fundamental and Transversal Skills Development, funded by the European Commission (grant number: 1010094837). The project aims to enhance reading skills within the adult population by creating an intelligent system that assesses text complexity and recommends suitable reading materials to adults with low literacy skills, contributing to reducing skills gaps and facilitating access to information and culture (https://iread4skills.com/). This dataset is the result of specifically devised classification and annotation tasks, in which selected texts were organized and distributed to trainers in Adult Learning (AL) and Vocational Education Training (VET) Centres, as well as to adult students in AL and VET centres. This task was conducted via the Qualtrics platform. The Dataset 2: annotated corpora by level of complexity for FR, PT and SP is derived from the iRead4Skills Dataset 1: corpora by level of complexity for FR, PT and SP ( https://doi.org/10.5281/zenodo.10055909), which comprises written texts of various genres and complexity levels. From this collection, a subset of texts was selected for classification and annotation. This classification and annotation task aimed to provide additional data and test sets for the complexity analysis systems for the three languages of the project: French, Portuguese, and Spanish. The texts in each of the language corpora were selected taking into account the diversity of topics/domains, genres, and the reading preferences of the target audience of the iRead4Skills project. This percentage amounted to the total of 462 texts per language, which were divided by level of complexity, resulting in the following distribution: · 140 Very Easy texts · 140 Easy texts · 140 Plain texts · 42 More Complex texts. Trainers were asked to classify the texts according to the complexity levels of the project, here informally defined as: Very Easy (everyone can understand the text or most of the text). Easy (a person with less than the 9th year of schooling can understand the text or most of the text) Plain (a person with the 9th year of schooling can understand the text the first time he/she reads it) More complex (a person with the 9th year of schooling cannot understand the text the first time he/she reads it). They were also asked to annotate the parts of the texts considered complex according to various type of features, at word-level and at sentence-level (e.g., word order, sentence composition, etc.), according to following categories: Lexical/word-related features - unknown word - word too technical/specialized or archaic - complex derived word - points to a previous reference that is not obvious - word (other) Syntactic/sentence-level features - unusual word order - too much embedded secondary information - too many connectors in the same sentence - sentence (other) - other (please specify) The sets were divided in three parts in Qualtrics and, in each part, the texts are shown randomly to the annotator. Students were asked to confirm that they could read without difficulty texts adequate to their literacy level. Each set contained texts from a given level, plus one text of the level immediately above. They were also asked to annotate words or sequences of words in the text that they did not understand, according to the following categories: - difficult word - difficult part of the text The complete results and datasets are in TSV/Excel format, in pairs of two files, with one file concerning the results from the classification (trainers)/validation (students) task and one file concerning the results from the annotation task. The complete datasets will be available under creative CC BY-NC-ND 4.0
iRead4Skills数据集2:面向法语(FR)、葡萄牙语(PT)和西班牙语(SP)的复杂度分级标注语料库,是一组按复杂度等级分类、并针对复杂度特征进行标注的文本集合,以XLSX格式存储。该语料库由iRead4Skills项目团队编译、分类并标注——该项目全称为「面向基础与通用技能发展的智能阅读提升系统」,由欧盟委员会资助(资助编号:1010094837)。项目旨在通过构建一套智能系统,评估文本复杂度并为低读写能力成年人推荐适配的阅读材料,以此提升成年人群的阅读技能,助力缩小技能差距,促进信息与文化的获取(详见项目官网:https://iread4skills.com/)。 本数据集源自专门设计的分类与标注任务:研究人员将筛选后的文本分发给成人学习(Adult Learning, AL)与职业教育与培训(Vocational Education and Training, VET)中心的培训师,以及该类中心的成年学员。本次任务通过Qualtrics平台完成。 iRead4Skills数据集2:面向法语、葡萄牙语和西班牙语的复杂度分级标注语料库,衍生自iRead4Skills数据集1:面向法语、葡萄牙语和西班牙语的复杂度分级语料库(https://doi.org/10.5281/zenodo.10055909),后者包含多种体裁与复杂度等级的书面文本。研究人员从该语料库中筛选出子集文本用于分类与标注任务,旨在为项目涉及的三门语言(法语、葡萄牙语、西班牙语)的复杂度分析系统提供额外的训练数据与测试集。 各语言语料库中的文本筛选兼顾了主题/领域、体裁的多样性,以及iRead4Skills项目目标受众的阅读偏好。每门语言最终筛选出共计462篇文本,按复杂度等级划分如下: · 140篇极简易文本 · 140篇简易文本 · 140篇普通文本 · 42篇较复杂文本 培训师需按照项目设定的复杂度等级对文本进行分类,项目对各等级的非正式定义如下: - 极简易:所有人均可理解该文本或大部分内容 - 简易:未完成九年级教育的人群可理解该文本或大部分内容 - 普通:完成九年级教育的人群可在首次阅读时理解该文本 - 较复杂:完成九年级教育的人群无法在首次阅读时理解该文本 此外,培训师需按词级与句级的多种特征类型,对文本中被认定为复杂的部分进行标注,标注类别包括: #### 词汇/词相关特征 - 陌生词汇 - 过于专业/小众或古旧的词汇 - 复杂派生词 - 指向不明显前置指代的词汇 - 其他词汇类问题 #### 句法/句级特征 - 非常规词序 - 嵌入过多次要信息 - 单句内连接词过多 - 其他句法类问题 - 其他(请注明) Qualtrics平台将本次任务分为三个部分,每部分内的文本将随机展示给标注者。 学员需确认可无障碍阅读适配其读写水平的文本。每个任务组包含对应等级的文本,以及一篇难度略高于该等级的文本。 此外,学员需按以下类别,对文本中无法理解的单词或词序列进行标注: - 难解词汇 - 难解文本片段 完整的实验结果与数据集以TSV/Excel格式存储,分为两组文件:一组对应分类(培训师)/验证(学员)任务结果,另一组对应标注任务结果。完整数据集将采用CC BY-NC-ND 4.0许可协议发布。



