Shanghainese Word Frequency (v1.0): a spoken-weighted frequency list with rank uncertainty
收藏资源简介:
A word frequency list for Shanghainese (Shanghai Wu, ISO 639-3 wuu): 2,524 ranked words, each with counts in conversational speech and in written Wu, a combined frequency per million (0.75 × spoken + 0.25 × written), document range, Gries' DP, and a bootstrap 95% rank interval. Limits. The spoken corpus is small: 20 speakers, about 4 hours, 52,412 tokens. Ranks are reliable only to rank 236, and there only to within about ±50%; at no depth are they within ±25%. Below rank 236 the order is indicative only. Split-half Spearman between random halves of the speakers is 0.80 (top 100), 0.77 (top 300) and 0.71 (top 500). Spoken and written counts agree weakly (ρ = 0.36 over 833 shared words), which is why speech is weighted 0.75. Transcription conventions affect some counts: 仔 and 哉 never occur in the conversations while 了 occurs 972 times. No native speaker has reviewed the list. Sources (counts only; no source text is redistributed). Spoken: MagicData, ASR-CShhiDiaCSC Chinese Shanghai Dialect Conversational Speech Corpus, via the Hugging Face copy TingChen-ppmc/Shanghai_Dialect_Conversational_Speech_Corpus, revision f10f366. Written: the Wu Chinese Wikipedia dump of 2026-09-01, filtered to colloquial Shanghai-type Wu: 2,405 of 48,393 articles, 459,388 tokens. Historical check, not part of the ranking: Pott (1907) and Edkins (1868) from Project Gutenberg, 34,534 tokens; 86% / 81% / 77% of the top 100 / 300 / 500 words appear there. Method. OpenCC conversion and a hand-written spelling map (the transcripts write 吾 for 我, and 伐 for both the question particle and the negator 勿); two context rules for 伐 and 呃 showed about 2–5% error in hand-checked samples of 80. Segmentation uses a 3,113-entry hand-written lexicon with a unigram model re-estimated by hard EM; particles such as 仔, 个 and 伐 are words and are never joined to a neighbour. Each word has a Wu Association romanisation, part of speech, flags and an English gloss written for this dataset. The release rebuilds byte for byte from pinned sources (seed 20260926). Files. The zip holds the list (output/shanghainese_word_frequency.tsv, also uploaded separately for preview), per-source counts, JSON files with every quoted number, the lexicon, the code and the documentation. The same release is on Figshare: https://doi.org/10.6084/m9.figshare.34001649 Licence. CC0 1.0: the release contains only counts, statistics and original lexicon, glosses and code. The source corpora keep their own licences (MagicData: CC BY-NC-ND 4.0; Wikipedia: CC BY-SA 4.0) and are not redistributed. Compiled by the Verbavia project. With thanks to MagicData, the contributors to the Wu Chinese Wikipedia, F. L. Hawks Pott and J. Edkins.



