遇见数据集

MonoKab – One Language, One Resource: A Monolingual Corpus for the Low-Resource Kabyle Language

收藏
Zenodo2026-09-26 更新2026-10-01 收录
官方服务:

资源简介:

Description of the Dataset MonoKab v1.0 is a large-scale monolingual Kabyle corpus containing 10,790,171 tokens across 1,768,427 sentences. The corpus is centered exclusively on Kabyle and is intended to support research in natural language processing (NLP), particularly research on low-resource languages (LRLs), language modeling, and the development and evaluation of computational resources for Kabyle. Data Sources and Processing The data were compiled from multiple external sources, including Hugging Face Datasets, the OPUS platform, GitHub repositories, publicly accessible websites, and other linguistic resources. A primary objective of MonoKab is to centralize existing Kabyle language data, particularly publicly accessible resources, into a single, organized corpus (That’s why: One Language, One Resource). By bringing together data that are otherwise distributed across multiple platforms and sources, MonoKab aims to make these resources easier to discover, access, process, and reuse for research and language technology development, while respecting the licensing and usage conditions of the original sources. The collected data underwent several processing steps, including cleaning, normalization, verification, and deduplication, before being integrated into MonoKab. The resulting corpus is organized at the text level, with each record containing a unique ID and its corresponding text. Metadata and source information are provided separately in an accompanying file, allowing the provenance of the data to be traced without duplicating this information in the main corpus file. Academic Context This work was conducted in an academic research context at the Higher National School of Computer Science (ESI), Algiers, and the National Conservatory of Arts and Crafts (CNAM), Paris, which served as the host institution. The project focuses on Natural Language Processing (NLP) for low-resource languages, with Tamazight as the broader research domain and Kabyle as a specific case study. Kabyle is a Northern Tamazight (Berber) language. The project aims to contribute to the digital inclusion of low-resource languages and support their integration into modern NLP systems. Licensing and Data Rights MonoKab v1.0 is provided under restricted access. Users granted access to the dataset must comply with the terms and conditions specified in the accompanying LICENSE.txt file. The corpus was compiled from external sources with heterogeneous licensing and usage conditions. MonoKab does not claim ownership of third-party materials included in the corpus. The inclusion of third-party material in the corpus does not modify or supersede the licensing terms applicable to the original source material. Users are responsible for reading and complying with LICENSE.txt and any applicable third-party licensing terms before accessing or using the dataset. Access Access to MonoKab v1.0 is subject to the applicable access conditions and licensing terms. Questions concerning access to the corpus or permitted uses should be directed to the MonoKab maintainers. Related Datasets ParaKab – Many Languages, One Kabyle: A Multilingual Parallel Corpus for a Low-Resource Language Version 1.110.5281/zenodo.22957880 Version 1.010.5281/zenodo.18875317

提供机构:
Zenodo
创建时间:
2026-09-25
二维码
社区交流群
二维码
科研交流群
商业服务