airudit/UD_v2_17_POS_LEMMA_GSD
收藏资源简介:
该数据集是一个用于自然语言处理任务的法语文本数据集,主要用于词性标注。它包含句子ID、词元列表、词根列表和通用词性标注(UPOS)等特征,其中UPOS涵盖17个类别,如名词、动词、形容词等。数据集分为训练集(14450个样本)、开发集(1476个样本)和三个测试集(分别基于不同来源:test_fr_gsd有416个样本,test_fr_pud有1000个样本,test_fr_poitevindivital有239个样本),适用于法语语言模型的训练和评估。
This dataset is a French text dataset for natural language processing tasks, primarily focused on part-of-speech tagging. It includes features such as sentence ID, token lists, lemma lists, and universal part-of-speech (UPOS) tags, with UPOS covering 17 categories like NOUN, VERB, and ADJ. The dataset is split into training set (14,450 examples), development set (1,476 examples), and three test sets (test_fr_gsd with 416 examples, test_fr_pud with 1,000 examples, and test_fr_poitevindivital with 239 examples), suitable for training and evaluating French language models.




