Mkulima Corpus: A Swahili Agricultural Corpus and Retrieval Benchmark for Domain-Specific NLP
收藏资源简介:
Mkulima (Swahili for "farmer") is a domain-specific Swahili agricultural corpus and retrieval test collection for natural language processing and information retrieval. The corpus comprises 4,021 documents totalling 35.2 million characters and 5.4 million words, drawn from five source types published between 2010 and 2026: government reports, FAO and NGO publications, agricultural extension materials, news articles and farmer-oriented blogs. Its vocabulary contains 137,980 word types, where a type is a lowercased token of two or more Latin letters or apostrophes. The corpus covers 25 agricultural topics including staple crops (mahindi, mpunga, muhogo), cash crops (kahawa, korosho, pamba), livestock (mifugo, kuku, ufugaji), fisheries (samaki, uvuvi) and agricultural inputs (mbolea, mbegu, umwagiliaji). Each document carries provenance metadata recording source type, origin, publication date and document category. CONTENTS corpus/mkulima_redistributable.jsonl — the 1,274 documents whose text may be redistributed (government, FAO/NGO, extension and blog material) reconstruction/news_manifest.jsonl — the 2,747 news documents as identifiers, source URLs, metadata and SHA-256 text hashes, with collection scripts benchmark/queries.jsonl — 50 Swahili agricultural retrieval queries benchmark/qrels.txt — 1,000 binary relevance judgments in 4-column TREC format benchmark/run_bm25.jsonl, benchmark/run_bm25_reranked.trec — retrieval runs benchmark/per_query_metrics.csv — per-query nDCG@10, MAP and MRR benchmark/duplicate_texts.csv — exact-duplicate document groups scripts/ — evaluation and significance-testing code MANIFEST.sha256 — checksums for every released file LICENSES.md, README.md — per-component licensing and documentation BENCHMARKS Domain-adaptive language modelling: fine-tuning AfroXLMR-base on Mkulima reduces perplexity on held-out agricultural Swahili text from 3.44 to 3.13, an 8.8% relative reduction. Agricultural document retrieval: the test collection provides 50 queries with 1,000 manually judged query-document pairs that completely cover the BM25 candidate pool to depth 20, with Cohen's kappa of 0.88 on a doubly annotated subset of 100 pairs. BM25 reaches nDCG@10 0.6041, MAP 0.5870 and MRR 0.7463. Rescoring the BM25 top-20 with the bge-reranker-v2-m3 cross-encoder reaches 0.6134, 0.5811 and 0.6896. Every reported figure is recomputed from the released artifacts by the released evaluation script. LICENSING Author-created materials (queries, judgments, runs, per-query results, metadata, topic taxonomy and all code) are released under CC BY 4.0. Document texts follow their publishers' terms, and news article text is not redistributed. See LICENSES.md for the per-component statement. CHANGES FROM VERSION 1 Version 1 distributed the full corpus, including news article text, under a single CC BY 4.0 record licence, and included a relevance-judgment file with additional title and snippet columns. Version 2 corrects both. News text is replaced by a reconstruction manifest, licensing is stated per component, the judgments are supplied as clean 4-column TREC, and the retrieval runs, per-query results, duplicate list, evaluation scripts and a checksum manifest are added. The corpus date range is corrected from 1995-2026 to 2010-2026 following an audit of document metadata and stated publication dates.



