somosnlp-hackathon-2022/readability-es-caes
收藏资源简介:
--- annotations_creators: - other language_creators: - other language: - es license: - cc-by-4.0 multilinguality: - monolingual size_categories: - unknown source_datasets: - original task_categories: - text-classification task_ids: [] pretty_name: readability-es-caes tags: - readability --- # Dataset Card for [readability-es-caes] ## Dataset Description ### Dataset Summary This dataset is a compilation of short articles from websites dedicated to learn Spanish as a second language. These articles have been compiled from the following sources: - [CAES corpus](http://galvan.usc.es/caes/) (Martínez et al., 2019): the "Corpus de Aprendices del Español" is a collection of texts produced by Spanish L2 learners from Spanish learning centers and universities. These text are produced by students of all levels (A1 to C1), with different backgrounds (11 native languages) and levels of experience. ### Languages Spanish ## Dataset Structure Texts are tokenized to create a paragraph-based dataset ### Data Fields The dataset is formatted as a json lines and includes the following fields: - **Category:** when available, this includes the level of this text according to the Common European Framework of Reference for Languages (CEFR). - **Level:** standardized readability level: simple or complex. - **Level-3:** standardized readability level: basic, intermediate or advanced. - **Text:** original text formatted into sentences. ## Additional Information ### Licensing Information https://creativecommons.org/licenses/by-nc-sa/4.0/ ### Citation Information Please cite this page to give credit to the authors :) ### Team - [Laura Vásquez-Rodríguez](https://lmvasque.github.io/) - [Pedro Cuenca](https://twitter.com/pcuenq) - [Sergio Morales](https://www.fireblend.com/) - [Fernando Alva-Manchego](https://feralvam.github.io/)
## 数据集元数据 - 注释创建者:其他 - 语言创建者:其他 - 语言:西班牙语(es) - 许可协议:CC BY 4.0 - 多语言属性:单语言 - 数据规模类别:未知 - 源数据集:原始数据集 - 任务类别:文本分类 - 任务子类别:无 - 数据集简称:readability-es-caes - 标签:可读性 --- # 「readability-es-caes」数据集卡片 ## 数据集说明 ### 数据集概览 本数据集收录自面向西班牙语作为第二语言教学的网站上的短篇文章,这些文章的来源如下: - [CAES语料库(CAES Corpus)](http://galvan.usc.es/caes/)(Martínez等,2019):《西班牙学习者语料库(Corpus de Aprendices del Español,简称CAES Corpus)》收录了来自西班牙语学习中心及高校的西班牙语第二语言学习者(L2 Learners)所撰写的文本。这些文本的作者覆盖全部水平等级(A1至C1),拥有11种不同母语背景,且学习经验各异。 ### 语言说明 西班牙语 ## 数据集结构 文本已完成分词处理,构建为基于段落的数据集。 ### 数据字段 本数据集采用行式JSON(JSON Lines)格式存储,包含以下字段: - **类别**:若可用,则包含依据《欧洲语言共同参考框架(Common European Framework of Reference for Languages,CEFR)》标注的文本等级。 - **等级**:标准化可读性等级:简易或复杂。 - **三级等级**:标准化可读性等级:基础、中级或高级。 - **文本**:已按分句格式整理的原始文本。 ## 附加信息 ### 许可信息 https://creativecommons.org/licenses/by-nc-sa/4.0/ ### 引用说明 请引用本数据集页面以标注作者贡献:) ### 开发团队 - [劳拉·巴斯克斯-罗德里格斯(Laura Vásquez-Rodríguez)](https://lmvasque.github.io/) - [佩德罗·昆卡(Pedro Cuenca)](https://twitter.com/pcuenq) - [塞尔吉奥·莫拉莱斯(Sergio Morales)](https://www.fireblend.com/) - [费尔南多·阿尔瓦-曼切戈(Fernando Alva-Manchego)](https://feralvam.github.io/)
数据集概述
数据集描述
数据集总结
本数据集是由专注于学习西班牙语的网站上的短篇文章汇编而成。这些文章主要来源于以下资源:
- CAES corpus (Martínez et al., 2019):“Corpus de Aprendices del Español”是一个由西班牙语作为第二语言学习者编写的文本集合,这些文本由来自学习中心和大学的学生编写,涵盖所有级别(A1至C1),具有不同的背景(11种母语)和经验水平。
语言
西班牙语
数据集结构
文本已分词,形成基于段落的数据集。
数据字段
数据集采用json lines格式,包含以下字段:
- Category: 根据欧洲共同框架(CEFR),当可用时,包括文本的级别。
- Level: 标准化可读性级别:简单或复杂。
- Level-3: 标准化可读性级别:基础、中级或高级。
- Text: 原始文本,格式化为句子。
附加信息
许可信息
本数据集遵循Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License。




