uyghur-headword-lists: headword inventories from three printed Uyghur dictionaries
收藏资源简介:
Open, machine-readable headword lists extracted from three printed Uyghur dictionaries: 130,218 written forms and 48,245 sub-entry phrases in the Uyghur Arabic script. Every entry is traceable to the scanned page column it was extracted from. Each also carries a per-round audit status recording whether a native speaker has verified it against the original scan; 16,561 forms (12.7%) are verified so far. Orthography follows the 2009 imla standard. Released as plain TSV under CC BY 4.0. Chinese transliterations tagged in the source dictionaries' own etymology columns were reviewed individually. 1,101 rows were removed, and the full removal list is published so the decision can be audited or reversed. The release also includes a benchmark quantifying frontier-model (gemini-3.1-pro-preview) reliability on post-OCR correction of Uyghur, now covering two dictionaries. Of 8,240 machine-proposed corrections verified by a native speaker, 71.4% were correct. Accuracy ranged from 100% on layout errors to near zero on judgement tasks. The model also showed a directional bias in rounded vowels: replacing ۇ with و was right 6% of the time, while the reverse was right 91%. This weakness replicated across two independently audited volumes. This dataset is corrected continuously. Cite the version DOI for reproducibility. I added the verified figure because that number grows the most between releases. I also replaced the "7% on dictionary structure" example with the vowel finding. The 7% was specific to the 2009 book, since the same class scored 92% in 2011, so it no longer describes the benchmark as a whole.



