Dasinya: A Badini Kurdish Corpus and BDPT Preprocessing Pipeline
收藏资源简介:
Dasinya is the first publicly available corpus of Badini Kurdish, a low-resource dialect written in Perso-Arabic script. The corpus comprises 107 validated source documents across five genres (books, news, broadcast, academic, and social media), containing 87,545 sentences and approximately 1.28 million words. The repository also includes the Badini Dialect Processing Toolkit (BDPT) — the first NLP preprocessing pipeline developed specifically for Badini Kurdish. The pipeline covers six stages: (1) file validation, (2) Unicode normalization with ZWNJ preservation, (3) structural cleaning, (4) deep cleaning, (5) text preprocessing, and (6) sentence segmentation. Stage 6 achieves a specialist-validated accuracy of 90.0% across all 107 files. Also included is BDPT_App_v3.html — a standalone browser-based interactive tool that runs the full six-stage pipeline without any installation required. This dataset is published in conjunction with the Data in Brief article: Saeed, V.A. & Jacksi, K. (2026). Dasinya: A Badini Kurdish Corpus and BDPT Preprocessing Pipeline.



