遇见数据集

Adyan: A High-Quality Automated NER Dataset for Sorani Kurdish

收藏
Mendeley Data2026-04-18 收录
官方服务:

资源简介:

The Named Entity Recognition Dataset for Kurdish Sorani (Adian) is an automated annotated resource created to support NLP studies for Kurdish Sorani, a low-resource language. By providing a large, structured dataset that trains and evaluates Kurdish language models, it addresses the lack of annotated corpora for nominal entity identification (NER). The dataset of more than 2,300 news articles published in 2024 by trusted Kurdish media sources such as Rudaw, NRT,Avanews and Kurdsatnews covers six main domains: politics, economy, sports, culture, interviews, and technology. It contains 654,404 tokens across 22,801 strings. A predefined dictionary of 12,030 named entities was used for annotation, covering 15 entity types. Annotation was performed automatically using a dictionary-based approach and the BIO tagging scheme. Preprocessing steps such as text cleaning, normalization, and duplicate removal were implemented using Python scripts to improve consistency and quality. Adyan is publicly available for academic and research use. Its scale, domain coverage and automated methodology make it a valuable resource not only for NER but also for sentiment analysis, tool translation and other Kurdish NLP tasks. The dataset makes a significant contribution to the advancement of language technology for underrepresented languages.

库尔德索拉尼语命名实体识别数据集(Adian)是专为低资源语言库尔德索拉尼语的自然语言处理(Natural Language Processing,NLP)研究打造的自动化标注资源。本数据集通过提供大规模结构化语料库以训练与评估库尔德语语言模型,填补了命名实体识别(Named Entity Recognition,NER)领域标注语料匮乏的空白。 该数据集收录了2024年由鲁道(Rudaw)、NRT、Avanews、Kurdsatnews等权威库尔德媒体发布的2300余篇新闻稿件,覆盖政治、经济、体育、文化、访谈、科技六大领域。语料库共包含22801个文本片段,累计654404个Token。标注环节采用包含12030个命名实体的预定义词典,覆盖15类实体类型,并通过基于词典的自动化方法与BIO标注方案完成标注。数据集借助Python脚本完成文本清洗、归一化、去重等预处理流程,以提升语料的一致性与质量。 Adyan可面向学术与科研用途公开获取。其规模、领域覆盖范围与自动化构建方法,使其不仅可用于命名实体识别任务,还可应用于情感分析、工具翻译等其他库尔德语自然语言处理任务。本数据集为弱势语言的语言技术发展作出了重要贡献。

创建时间:
2026-02-26
二维码
社区交流群
二维码
科研交流群
商业服务