A Corpus of Asante Twi
收藏资源简介:
This is a ~14 million tokens preprocessed corpus of Asante Twi, realized for the Luminos Fund's Bilingual Boost initiative thanks to funding from the Gates Foundation. An interactive, searchable and visualization-rich version of this corpus is vailable at https://asante-twi-corpus.fly.dev/app/ Corpus composition Subcorpus Documents Tokens Share Doctrinal 3,214 12,125,333 86.5% General 2,246 1,353,049 9.7% Children 190 535,293 3.8% Total 5,650 14,013,675 100% The corpus also includes 34,647 Twi-English aligned sentences Data Quatlity Preprocessing accuracy has been evalauted against a single-annotator triple-proofread set of 13 documents for a total of 12,162 proofed tokens (10,372 words + punctuation; 9,460 development / 2,702 sealed held-out, expert-proofed and re-tokenized to the delivered corpus). Accuracy, words-only (punctuation excluded — it is tagged trivially and inflates scores): Metric · words-only (excl. punctuation) Dev · 8,028 words Held-out · 2,344 words Lemma accuracy (row · aligned micro-F1) 0.915 0.902 — correct / missed 7,342 / 686 2,115 / 229 Lemma F1 (bag-of-tokens) 0.926 0.908 PoS accuracy (token) 0.873 0.827 PoS macro-F1 (per-tag) 0.799 0.788 — all-token (incl. punctuation): lemma / PoS acc 0.927 / 0.892 0.914 / 0.846 Lemma accuracy (~0.91) and PoS macro-F1 (~0.79) hold across development and the sealed held-out set, so the combined 12,162-token gold standard confirms the pipeline generalizes rather than overfits. 5-fold spread (words-only): dev lemma 0.915 ± 0.005 / PoS 0.873 ± 0.008; held-out lemma 0.902 ± 0.011 / PoS 0.827 ± 0.024. Resources Included 5,650 corpus files one metadata file detailing genre, title, author... of each corpus document 34,647 Twi-English aligned sentences



