SAMER阿拉伯语文本简化语料库
收藏资源简介:
SAMER阿拉伯语文本简化语料库是由纽约大学阿布扎比分校计算语言学建模实验室创建的,旨在为学龄学习者提供文本简化资源。该数据集包含从15本公开可用的阿拉伯语小说中选取的约15.9万字文本,这些小说大多在1865年至1955年间出版。数据集不仅包括文档和单词级别的可读性标注,还为每个文本提供了两种简化版本,针对不同可读性水平的学习者。创建过程中遵循严格的指导原则以确保标注质量。该数据集的应用领域包括阿拉伯语文本简化研究、自动可读性评估以及阿拉伯语教学语言技术的开发。
SAMER Arabic Text Simplification Corpus was developed by the Computational Linguistics Modeling Lab at New York University Abu Dhabi, aiming to provide text simplification resources for school-age learners. This corpus contains approximately 159,000 words of text selected from 15 publicly available Arabic novels, most of which were published between 1865 and 1955. In addition to document-level and word-level readability annotations, the dataset also provides two simplified versions for each text, tailored to learners at different readability levels. Strict guiding principles were followed throughout the creation process to ensure the quality of annotations. The potential application areas of this corpus include Arabic text simplification research, automatic readability assessment, and the development of language technologies for Arabic language teaching.




