Introduction This file contains documentation on the Buckwalter Arabic Morphological Analyzer Version 2.0. Data The data consists primarily of three Arabic-E
该数据集名为Bosque Part of Speech PT-PT,主要用于葡萄牙语的词性标注任务。数据集包含tokens、lemmas和pos_tags三个特征,均为字符串序列。数据集分为训练集和测试集,分别包含9071和576个示例。数据集的任务类别是token-classification,语言为葡萄牙语(pt),标签包括pos、pos-tagging和part-of-speech。