遇见数据集

Darbest Dataset: Universal Dependencies Treebank for Standard Sorani Kurdish

收藏
Mendeley Data2026-09-08 收录
官方服务:

资源简介:

The Darbest dataset is a Universal Dependencies (UD) treebank dataset for Standard Sorani Kurdish written in the Perso-Arabic script. It contains 69,000 annotated sentences and 1,205,855 tokens collected from nine textual domains. The corpus was collected from seven Kurdish online news websites and supplemented with texts from published books. Before preprocessing, the collected corpus contained 1,250,275 words from 5,627 web pages together with book-based texts and was preprocessed using a Python-based pipeline involving text cleaning, Unicode and punctuation normalization, sentence segmentation, and tokenization. The dataset was developed to provide a large-scale syntactically and morphologically annotated resource for Standard Sorani Kurdish. A separate 100-sentence gold-standard set was manually annotated according to the Universal Dependencies v2 guidelines. Sorani Kurdish linguistic experts supported the selection of sentences representing diverse and linguistically complex structures and reviewed LLM-generated annotations for errors. The gold-standard set was used to construct few-shot prompts for annotating the remaining corpus. The resulting annotations were represented in the standard CoNLL-U format and validated using the official Universal Dependencies validation tool, followed by manual correction and quality review. The released treebank is divided into training, development, and test sets and includes lemmas, Universal Part-of-Speech (UPOS) tags, morphological features, syntactic heads, and dependency relations. The dataset can be used to train, evaluate, and benchmark NLP models for part-of-speech tagging, lemmatization, morphological analysis, dependency parsing, and related computational linguistics tasks. The accompanying repository also contains the separate 100-sentence gold-standard set, plain-text corpus splits, README documentation, and a dataset statistics spreadsheet.

Darbest数据集是一款采用波斯阿拉伯字母书写的标准索拉尼库尔德语通用依存树库(Universal Dependencies, UD)数据集。其包含来自9个文本领域的69000条标注语句与1205855个词元(Token)。该语料库采集自7个库尔德语在线新闻网站,并辅以已出版书籍中的文本。预处理前,采集得到的语料库包含来自5627个网页与书籍文本的1250275个词汇,随后通过基于Python的流水线完成预处理,流程涵盖文本清洗、Unicode与标点符号归一化、语句分句与词元分词。 本数据集的开发旨在为标准索拉尼库尔德语提供大规模的句法与形态标注资源。研究人员依据通用依存标注体系v2指南,手动标注了100条语句的独立金标准集。索拉尼库尔德语语言专家参与了涵盖多样且语言结构复杂的语句筛选工作,并对大语言模型(Large Language Model, LLM)生成的标注结果进行了错误校验。该金标准集被用于构建少样本(Few-shot)提示词,以标注剩余语料库。最终得到的标注结果以标准CoNLL-U格式存储,并通过官方通用依存验证工具进行校验,随后经人工修正与质量评审。 本次发布的树库划分为训练集、开发集与测试集,涵盖词形还原结果、通用词类标注(Universal Part-of-Speech, UPOS)标签、形态特征、句法中心词与依存关系。该数据集可用于训练、评估与基准测试面向词性标注、词形还原、形态分析、依存句法分析及相关计算语言学任务的自然语言处理模型。配套的代码仓库还包含该100条语句的独立金标准集、纯文本语料拆分文件、README说明文档与数据集统计表格。

创建时间:
2026-08-17
二维码
社区交流群
二维码
科研交流群
商业服务