遇见数据集

Dasinya: A Badini Kurdish Corpus and BDPT Preprocessing Pipeline

收藏
Zenodo2026-06-17 更新2026-06-21 收录
官方服务:

资源简介:

Dasinya is the first publicly available corpus of Badini Kurdish, a low-resource dialect written in Perso-Arabic script. The corpus comprises 107 validated source documents across five genres (books, news, broadcast, academic, and social media), containing 87,545 sentences and approximately 1.28 million words. The repository also includes the Badini Dialect Processing Toolkit (BDPT) — the first NLP preprocessing pipeline developed specifically for Badini Kurdish. The pipeline covers six stages: (1) file validation, (2) Unicode normalization with ZWNJ preservation, (3) structural cleaning, (4) deep cleaning, (5) text preprocessing, and (6) sentence segmentation. Stage 6 achieves a specialist-validated accuracy of 90.0% across all 107 files. Also included is BDPT_App_v3.html — a standalone browser-based interactive tool that runs the full six-stage pipeline without any installation required. This dataset is published in conjunction with the Data in Brief article: Saeed, V.A. & Jacksi, K. (2026). Dasinya: A Badini Kurdish Corpus and BDPT Preprocessing Pipeline.

提供机构:
Zenodo
创建时间:
2026-06-17
二维码
社区交流群
二维码
科研交流群
商业服务