ParaKab – Many Languages, One Kabyle: A Multilingual Parallel Corpus for a Low-Resource Language
收藏资源简介:
Description of the Dataset ParaKab v1.1 is a multilingual, Kabyle-centric parallel corpus composed of five language pairs: Kabyle–English (72.87 %) Kabyle–French (25.93 %) Kabyle–Spanish (1.01 %) Kabyle–German (0.04 %) Kabyle–Arabic (0.16 %) Together, these resources contain 2,801,431 aligned sentence pairs, providing a large multilingual parallel dataset centered on Kabyle. The corpus is intended to support research in natural language processing (NLP), particularly for low-resource languages (LRLs), machine translation, multilingual NLP, cross-linguistic studies, and the development and evaluation of computational resources for Kabyle. Data Sources and Processing The data were compiled from multiple external sources, including Hugging Face Datasets, the OPUS platform, publicly accessible websites, and other linguistic resources. The collected data underwent processing including cleaning, normalization, alignment, verification, and deduplication before being integrated into ParaKab. The resulting corpus is organized by language pair and includes metadata describing the provenance and source information. The dataset records contain the following main fields: id — unique identifier assigned to the corpus record; lang — language-pair identifier; source — source-language sentence; target — Kabyle sentence; Academic Context This work was conducted in an academic research context at the Higher National School of Computer Science (ESI), Algiers, and the National Conservatory of Arts and Crafts (CNAM), Paris, which served as the host institution. The project focuses on Natural Language Processing (NLP) for low-resource languages, with Tamazight as the broader research domain and Kabyle as a specific case study. Kabyle is a Northern Tamazight (Berber) language. The project aims to contribute to the digital inclusion of low-resource languages and support their integration into modern NLP systems. Licensing and Data Rights ParaKab v1.1 is provided under restricted access. Users granted access must comply with the terms and conditions specified in the accompanying LICENSE.txt file. The dataset was compiled from external sources with heterogeneous licensing and usage conditions. ParaKab does not claim ownership of third-party materials included in the corpus. Users are responsible for reading and complying with LICENSE.txt before accessing or using the dataset. Access Questions concerning access to the corpus or permitted use should be directed to the ParaKab maintainers. Previous Version An earlier version of ParaKab (v1.0) is available under open access. Users seeking an openly accessible version of the corpus can refer to the ParaKab v1.0 dataset record. Related Datasets MonoKab – One Language, One Resource: A Monolingual Corpus for the Low-Resource Kabyle Language Version 1.010.5281/zenodo.22959239



