遇见数据集

CKL Corpus: A Manually Annotated Central Kurdish Corpus for Part-of-Speech Tagging

收藏
Zenodo2026-07-13 更新2026-08-01 收录
官方服务:

资源简介:

The CKL Corpus is a manually annotated part-of-speech corpus developed for Central Kurdish, also known as Sorani Kurdish. Version 1.0.0 contains 144,173 annotated tokens distributed across 3,612 sentences, with 20,717 unique word types. The corpus was developed from 20 Central Kurdish academic articles. The texts were normalized, segmented, tokenized, and manually annotated. Sentences identified as Arabic, Persian, or English were excluded. Each token may contain up to three parallel POS annotations (Tag1, Tag2, and Tag3) to represent linguistic ambiguity and morpho-syntactic information. Approximately 30.3% of the corpus tokens contain more than one valid annotation. The corpus uses an 86-tag fine-grained POS scheme. This scheme is a corpus-attested subset of the broader 97-tag Central Kurdish POS standard introduced by Sabr et al. (2025); tags that did not occur in the corpus were not instantiated. The deposited package contains the corpus data files, a UTF-8 machine-readable tagset file (CKL_POS_Tagset_86.tsv), a README, and a license notice. The CKL Corpus supports research in Central Kurdish natural language processing, including part-of-speech tagging, morphological analysis, sequence labelling, linguistic analysis, and the training and evaluation of machine-learning and deep-learning models. The corpus annotations and original documentation are released under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License (CC BY-NC-SA 4.0), subject to the third-party rights notice included in LICENSE.txt. Noncommercial sharing and adaptation are permitted with attribution and ShareAlike. Commercial use requires separate written permission from the corpus creators. Dataset DOI:https://doi.org/10.5281/zenodo.21320625 Associated corpus publication:Al-Raghefy, H., Maghdid, H. S., and Taher, A. H. (2026). A Novel Central Kurdish Part-of-Speech Corpus and Deep Tagging Model Evaluation. ARO—The Scientific Journal of Koya University, 14(1), 384–394. https://doi.org/10.14500/aro.12641 Source POS tagset publication:Sabr, S. S., et al. (2025). A Comprehensive Part-of-Speech Tagging to Standardize Central-Kurdish Language: A Research Guide for Kurdish Natural Language Processing Tasks. Journal of Studies in Science and Engineering, 5(2), 15–38. https://doi.org/10.53898/josse2025531 Users must cite the official Zenodo dataset record when using the corpus. For academic use of the complete corpus, users are strongly requested to also cite the ARO corpus paper and the JOSSE source-tagset paper. Users who reproduce, adapt, analyse, or discuss the POS tagset should cite the JOSSE paper. Contact information: For questions about the CKL Corpus, access requests, annotation corrections,or commercial-use permission, please contact: Haneen Al-RaghefyDepartment of Software Engineering, Faculty of Engineering, Koya UniversityEmail: haneen.hayder@koyauniversity.orgORCID: https://orcid.org/0009-0007-1867-6614

提供机构:
Zenodo
创建时间:
2026-07-13
二维码
社区交流群
二维码
科研交流群
商业服务