相关数据集
Buckwalter Arabic Morphological Analyzer Version 2.0
Introduction This file contains documentation on the Buckwalter Arabic Morphological Analyzer Version 2.0. Data The data consists primarily of three Arabic-E
DataCite Commons2021-07-01 更新220
Noun Verb Dataset
该数据集包含自然发生的英语句子,这些句子具有非平凡的名词-动词歧义。数据集用于帮助英语词性标注器改进在名词-动词歧义上的表现,从而提高翻译和文本到语音合成的下游任务的准确性。
github2024-02-26 更新140
liaad/Bosque_PT-PT
该数据集名为Bosque Part of Speech PT-PT,主要用于葡萄牙语的词性标注任务。数据集包含tokens、lemmas和pos_tags三个特征,均为字符串序列。数据集分为训练集和测试集,分别包含9071和576个示例。数据集的任务类别是token-classification,语言为葡萄牙语(pt),标签包括pos、pos-tagging和part-of-speech。
Hugging Face2023-12-12 更新60
prakod/GLuecos_POS_EN_HI_FG
--- dataset_info: features: - name: words sequence: string - name: label1 sequence: string - name: label2 sequence: string splits: - name: dev_Romanized num_bytes: 110280
Hugging Face2024-05-31 更新160
prakod/POS_hinglish_codemixed_tweets
--- dataset_info: features: - name: words sequence: string - name: label1 sequence: string - name: label2 sequence: string splits: - name: train num_bytes: 764955 num_e
Hugging Face2024-06-05 更新90



