CTAB: Corpus of Tunisian Arabizi
收藏资源简介:
This dataset has been created between 2017 and 2021 to provide a textual resource that can be used to study the behaviors of Tunisian people in writing Tunisian Arabic (ISO 693-3: aeb) in Latin Script. This corpus is constituted from messages written using Tunisian Arabic Chat Alphabet or Arabizi and is developed to solve the matter of the lack of NLP databases about the use of the Latin Script for transcribing Tunisian Arabic. The messages are automatically pulled using web scraping of Facebook public pages and are kept as they are without any annotation, spelling adjustments or morphological and syntactic labeling. Then, messages that are written in Latin Script but not in Tunisian Arabic are manually eliminated. Finally, every collection of messages that are retrieved from the same Facebook page in the same period is included in the same text file where every message is featured as one line.
本数据集于2017年至2021年间构建,旨在提供可用于研究突尼斯民众使用拉丁字母书写突尼斯阿拉伯语(ISO 693-3: aeb)时的书写行为的文本资源。该语料库由使用突尼斯阿拉伯语聊天字母(Arabizi)书写的消息构成,旨在解决当前突尼斯阿拉伯语拉丁转写领域自然语言处理(Natural Language Processing, NLP)数据库匮乏的问题。这些消息通过网页爬虫抓取Facebook公开页面自动获取,原始内容未经过任何标注、拼写调整或形态句法标注处理,直接予以保留。随后,人工剔除了使用拉丁字母书写但并非突尼斯阿拉伯语的消息。最终,同一时期内从同一Facebook页面抓取的所有消息集合将被存入同一个文本文件,其中每条消息单独占一行。



