官方服务:
资源简介:
64 Part-of-Speech-getaggte Dramen Calderón de la Barcas, txt-Dateien, UTF-8-codiert
应用场景:
相关数据集
Buckwalter Arabic Morphological Analyzer Version 2.0
Introduction This file contains documentation on the Buckwalter Arabic Morphological Analyzer Version 2.0. Data The data consists primarily of three Arabic-E
DataCite Commons2021-07-01 更新220
thongnef/MSB_preprocess
--- dataset_info: features: - name: sentence_idx dtype: int64 - name: word sequence: string - name: pos sequence: int64 - name: tag sequence: int64 splits: - name: train
Hugging Face2024-01-02 更新130
liaad/Bosque_PT-PT
该数据集名为Bosque Part of Speech PT-PT,主要用于葡萄牙语的词性标注任务。数据集包含tokens、lemmas和pos_tags三个特征,均为字符串序列。数据集分为训练集和测试集,分别包含9071和576个示例。数据集的任务类别是token-classification,语言为葡萄牙语(pt),标签包括pos、pos-tagging和part-of-speech。
Hugging Face2023-12-12 更新60
prakod/GLuecos_POS_EN_HI_FG
--- dataset_info: features: - name: words sequence: string - name: label1 sequence: string - name: label2 sequence: string splits: - name: dev_Romanized num_bytes: 110280
Hugging Face2024-05-31 更新160
Beseda Corpus Lemmatisation Lexicon
This lexicon contains inflected open class words from the [Dictionary of Standard Slovenian](http://bos.zrc-sazu.si/sskj_en.html) that are augmented by wordforms, their part of speech tags and their l
SSH Open MarketPlace2023-10-17 更新120



