遇见数据集

The Orthographic Vowel Rhythm of English: A Length-Stratified Positional Analysis

收藏
Zenodo2026-06-08 更新2026-06-12 收录
官方服务:

资源简介:

The Orthographic Vowel Rhythm of English: A Length-Stratified Positional Analysis Boicho Dimitrov Temelakiev · Saxon Ventura Research Ltd. *Working draft — a descriptive data report. Junction Grammar program; companion to "The Orthographic Junctions of English" (OJE).* 08th of June, 2026 CC BY --- Abstract We measure the per-position proportion of vowel letters ("vowel mass") in the English orthographic lexicon, stratified by word length N. On a deposited corpus of 455,246 word types (4,254,354 letters; mean length 9.345), three structures recur at every length from N=4 to N=16: consonant brackets at both word edges; a front vowel nucleus fixed at position 2; and a back vowel nucleus fixed at position N−2. As N grows, the front nucleus decays to a plateau while the back nucleus holds and then strengthens, producing a monotonic "handover" from front- to back-dominance that crosses parity near N≈13. To separate this positional structure from the universal consonant–vowel alternation of English, we compare the observed profiles against a two-state (vowel/consonant) Markov chain estimated from the same corpus. The chain — which is, by construction, the model A. A. Markov introduced in 1913 — relaxes to its stationary vowel rate within a few positions and cannot reproduce the observed profiles. The residual (observed minus Markov) is a low-amplitude interior pattern that is highly reproducible within a length cohort (split-half r ≈ 0.99) yet distinct between cohorts, so it functions as a length-specific signature rather than a single universal template. Two controls confirm the structure is carried by letter *order*: within-word shuffling collapses every cohort's positional range, and the two-letter (N=2) cohort, whose composition is near-uniform, yields a flat profile. This is a descriptive report: we present the measurements and the deposited corpus that reproduces them, and make no claim of novelty or priority. A section on related work points to the literature this sits among; where our measurements overlap existing findings, that overlap is unintentional and the prior work takes precedence. --- 1. Introduction That vowels and consonants are unevenly distributed across positions in English words is an old and widely noted fact, visible in everything from cryptanalytic letter-frequency tables to recreational word-game statistics. What is less often studied is how that positional structure *changes with word length*, and whether the structure that remains after the language's baseline consonant–vowel alternation is removed is stable enough to characterise a word by its length. This paper does two things, and is deliberately modest about scope. First, it reports the per-position vowel mass for every word-length cohort from N=2 to N=16, assembled into a single developmental table. Second, it removes the universal alternation backbone using a two-state Markov null and examines the residual. We make no claim that any of this is new. The measurements may overlap work we are unaware of; the corpus is deposited so that anyone can check them, and a closing section points to the related literature we do know of. We work in orthography, not phonology: the units are letters A–Z, the vowels are A, E, I, O, U, and Y is treated as a consonant throughout. We measure vowel *mass* — the fraction of cohort words carrying a vowel at a given position — rather than any positional mean, because averaging letter values across a categorical alphabet conflates distinct distributions. The analysis is confined to within-word letter adjacency in a type list; we make no claims about running text or token frequency. 2. Corpus and Method 2.1 Corpus and provenance The corpus derives from the public `words.txt` word list (dwyl). Entries containing non-alphabetic characters were removed, and the remainder was case-folded and deduplicated, yielding **455,246 unique alphabetic entries** comprising **4,254,354 letter occurrences** at a mean length of **9.345**. (The raw source contained roughly 11.3k more entries before cleaning.) Single-letter entries (n=25) are retained in the corpus but excluded from the per-position analysis, which requires at least two positions. Because the upstream list drifts over time, the exact 455,246entry file analysed here is included as a data file accompanying this Zenodo deposit and is the corpus of record; all figures are reproducible from it. A reader downloading the live upstream list today will not recover the same entry count. 2.2 Conventions Vowels: A, E, I, O, U. Y: consonant. For a length-N cohort, vowel mass at position p is the fraction of the cohort's words whose p-th letter is a vowel, reported as a percentage. The corpus-wide vowel rate is 38.9%, which serves as the baseline reference throughout. 2.3 The Markov null We estimate a two-state (V/C) first-order Markov chain on within-word letter adjacencies across the whole corpus: | transition | count | probability | |---|---:|---| | V → V | 205,064 | P(V\|V) = 0.132 | | V → C | 1,345,648 | P(C\|V) = 0.868 | | C → V | 1,348,038 | P(V\|C) = 0.600 | | C → C | 900,358 | P(C\|C) = 0.400 | The global vowel rate is 0.3885 and the chain's stationary vowel probability is π_V = 0.409. Seeded at a cohort's observed position-1 vowel rate, the chain is propagated forward to length N to give a position-by-position baseline; the **residual** is observed minus baseline, in percentage points. This is exactly the vowel/consonant Markov chain Markov (1913) introduced (see §7); our use of it is as a null to be subtracted, not as a model of the data. 2.4 Length cohorts and cutoff We analyse cohorts N=2 through N=16, which together are 98.48% of the corpus. The cutoff is set by statistical resolution, not convenience: cohort size falls through a ~5,000-word floor (below which positional texture is no longer resolved cleanly in this corpus) immediately after N=16 (N=16: n=6,173; N=17: n=3,505), and the long tail beyond is both underpowered and morphologically narrow (dominated by a single Latinate suffix family). N=2 is retained as a negative control. 2.5 Controls Two controls test whether the signal depends on letter order. (i) *Within-word shuffle*: letters of each word are permuted, preserving length and composition but destroying order; positional structure should vanish. (ii) *Split-half stability*: each cohort is split at random and the interior residual computed on each half; a length-specific signature should reproduce within a cohort far better than it matches other cohorts. 3. Results: the developmental spine Per-position vowel mass (%), N=2–16. Position 1 is the wordinitial letter; position N the final letter. N | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 ---+---------------------------------------------------------------- 2 | 21 22 3 | 23 38 25 4 | 19 62 34 33 5 | 17 64 34 44 33 6 | 18 66 30 38 54 29 7 | 18 64 32 32 50 48 25 8 | 19 61 33 37 36 53 46 23 9 | 22 58 32 36 42 36 54 45 23 10 | 24 56 31 38 39 43 37 51 48 22 11 | 24 55 32 37 40 36 48 36 50 50 20 12 | 26 52 32 38 39 38 40 47 37 51 50 18 13 | 26 53 31 38 41 36 44 37 49 36 53 49 17 14 | 27 52 30 38 40 38 40 41 40 49 35 56 46 16 15 | 28 52 30 39 39 38 43 38 43 41 50 35 58 44 13 16 | 29 50 30 41 38 37 42 38 41 44 41 51 35 58 40 13 Four regularities hold across the whole table. **Consonant brackets at both edges.** Position 1 is well below the 38.9% baseline at every length and *rises* monotonically with N (17% at N=5 → 29% at N=16), consistent with the vowelinitial Latinate prefixes (un-, in-, over-) that longer words admit. The final position *deepens* monotonically (33% at N=4 → 13% at N=16): the hardest word-final consonant sits in the longest words. **A front nucleus fixed at position 2.** Position 2 peaks at 66% (N=6), then decays toward a plateau of ~50–52% for long words. Its emergence is sharp between N=3 (position 2 = 38%, near baseline) and N=4 (62%, full nucleus); at N=3 the profile is transitional, with high-frequency function-word onsets (the, she, why) pulling position 2 down. **A back nucleus fixed at position N−2.** The penultimate-butone position carries a vowel nucleus that holds at ~50–54% and then strengthens to 58% in the longest cohorts — the suffix vowel of the -IZATION / -IFICATION family, where the single strongest letter spike in the table (I, ~28% at N=16, position 14) occurs. **N=2 negative control.** The two-letter cohort is nearuniform in composition (it enumerates admissible letter pairs rather than sampling vocabulary), and its profile is correspondingly flat (21, 22). A flat input yielding a flat profile confirms the method does not manufacture consonantvowel structure. 4. Results: the handover Defining the front-minus-back gap as (position-2 mass) − (back-nucleus mass), the gap declines monotonically with length and changes sign: | N | 4 | 6 | 8 | 10 | 12 | 13 | 14 | 15 | 16 | |---|---|---|---|---|---|---|---|---|---| | front−back gap (pts) | +29 | +12 | +9 | +6 | +1 | −0.7 | −4.7 | −6.0 | −8.2 | The crossover near N≈13 is not a positional collision — the two nuclei stay anchored at position 2 and at N−2 throughout — but an *amplitude* crossover: the front nucleus decays to its plateau while the back nucleus holds and then rises. Short words are front-weighted; long words are back-weighted. 5. Results: the interior signature Subtracting the Markov baseline isolates the part of the profile the universal alternation cannot explain. The interior residual (observed − Markov, percentage points) for positions 2…N−1, front-aligned: N | p2 p3 p4 p5 p6 p7 p8 p9 p10 p11 p12 p13 p14 p15 ----+----------------------------------------------------------------------------------- 4 | +11.2 -2.4 5 | +12.1 -1.6 +0.3 6 | +14.1 -6.0 -4.7 +14.2 7 | +12.7 -4.2 -11.0 +10.7 +7.0 8 | +10.2 -3.5 -5.8 -3.9 +11.2 +5.9 9 | +8.6 -4.4 -6.7 +2.3 -5.7 +13.5 +4.1 10 | +7.6 -6.0 -5.0 -1.1 +1.5 -3.3 +9.7 +7.4 11 | +6.1 -5.7 -5.2 -0.0 -5.3 +7.6 -5.0 +9.6 +8.8 12 | +4.3 -5.7 -3.9 -1.2 -3.0 -0.4 +5.7 -4.1 +10.3 +9.1 13 | +4.8 -7.1 -4.3 +0.7 -4.7 +3.0 -3.5 +8.2 -4.5 +12.5 +8.4 14 | +4.4 -7.3 -3.8 +0.2 -2.7 -0.6 +0.2 -1.1 +8.4 -5.6 +15.5 +5.2 15 | +5.0 -8.1 -3.1 -1.2 -2.7 +2.0 -3.1 +2.4 +0.3 +8.8 -6.1 +17.1 +3.2 16 | +3.6 -7.9 -0.9 -2.1 -4.0 +1.1 -3.2 -0.3 +3.3 +0.0 +9.6 -5.7 +17.5 -0.4 The same grammar repeats at every length: a positive head at position 2 (the front-nucleus excess over the chain, decaying with N as the chain catches up); a reliable negative trough at position 3; a large positive spike at position N−2 (the back nucleus) that grows in amplitude and steps one column right per unit of N; and, between them, a low-amplitude valley that lengthens and flattens. The residual amplitude is concentrated at the two ends of the interior and relaxes toward zero in the middle — the same edge concentration seen in the whole word, one level down. **Reproducibility of the residual.** Within a cohort, splithalf interior residuals correlate at r ≈ 0.99 (for cohorts with ≥4 interior points). Between cohorts, residuals correlate far less (front-aligned ≈ 0.63; back-aligned ≈ 0.51). The within-cohort stability greatly exceeding between-cohort similarity is what licenses calling the residual a *lengthspecific* signature rather than a universal shape: it is a fingerprint indexed by N. 6. Results: controls **Shuffle.** Within-word permutation collapses each cohort's positional range from 35–45 points to under 2 points. The statistic therefore reads letter *order*, not mere composition. **Markov relaxation.** The baseline chain reaches its stationary rate (π_V ≈ 0.41) within roughly five positions and stays there; it cannot produce the sustained interior zigzag or the anchored nuclei. The structure in the data is, by construction, what the chain leaves unexplained. 7. Related work we are aware of We make no claim of novelty or priority for anything in this report. This section points to the literature this work sits among, so a reader can place it; where our measurements overlap any of it, the overlap is unintentional and the prior work takes precedence. The list is what we happen to know of, not an exhaustive survey. **Two-state V/C Markov chains (Markov 1913; Shannon 1948, 1951).** The first application of Markov chains was Markov's hand count of vowel/consonant transitions in *Eugene Onegin*, which produced transition and stationary probabilities very close to ours on a different language and substrate. Our null model is that model; Shannon's letter n-gram work generalised it. We claim no novelty for the null itself — its orthodoxy is precisely why we use it as a baseline. Contemporary work continues in this vein: Sabatini (2026a, 2026b) applies a V/C Markov encoding to *running text* — Dante's *Commedia* and Pushkin's *Eugene Onegin* — to index graphemic dependency structure across a work, where transitions cross word boundaries and carry token weighting. The present analysis differs in substrate from the outset: it confines the chain to within-word adjacency over a deduplicated type list and uses it only as a positional null over length-stratified cohorts. The shared element is the two-state encoding; the objects measured (a text's dependency drift versus a lexicon's positional profile by length) are not comparable. **Speech-rhythm metrics (Ramus, Nespor & Mehler 1999; Grabe & Low 2002; Hirst 2009).** The %V family of metrics quantifies the proportion of vocalic material, but it is acoustic, computed over whole utterances, and explicitly disregards word boundaries and internal position. Hirst (2009) observed that %V largely tracks vowel proportion at the text level. Our quantity is a deliberate repurposing of the *idea* of vowel proportion as a positional, within-word, length-stratified orthographic measure; it is not the acoustic metric and should not be read as a rhythm-class typology. **Information front-loading and word-edge effects (Pimentel, Cotterell & Roark 2021; and the psycholinguistic wordrecognition literature).** There is cross-linguistic evidence that words front-load information, and explicit discussion of whether incremental recognition can *manufacture* apparent front-loading — the same kind of confound our shuffle control and Markov null happen to address. Our edge observations sit alongside this line, expressed in vowel mass rather than segment surprisal. **Successor-variety and morpheme-boundary work (Harris 1955; Hafer & Weiss 1974).** Cited in the companion OJE paper; relevant as the tradition of reading morphological structure off letter-transition statistics, which our boundary observations touch but do not directly extend here. 8. What this report contains, and its limits The table separates the measurements we are presenting from the background they rest on and the limits we place on them. No row asserts that a measurement is new. | # | Statement | Kind | |---|---|---| | 1 | Vowel mass varies by position; edges are consonantheavy, position 2 vowel-heavy | Background — long noted in positional letter statistics | | 2 | A two-state V/C Markov chain on English letters gives P(V\|V)≈0.13, P(V\|C)≈0.60, π_V≈0.41 | Background — Markov (1913) lineage; our values reproduce it on a new substrate | | 3 | Per-position vowel mass, computed *as mass not means*, *stratified by length N=2–16*, assembled into one developmental table | What we measured | | 4 | Front nucleus at position 2; back nucleus at position N−2; monotonic front→back handover crossing near N≈13 | What we measured | | 5 | The Markov residual is within-cohort-stable and betweencohort-distinct — a length-indexed interior pattern | What we measured | | 6 | Structure is order-dependent (shuffle collapses it) and not method-manufactured (N=2 control) | Method check | | 7 | All findings pertain to orthographic word *types* in within-word adjacency only | Scope limit — no claim about phonology, running text, or token frequency | | 8 | Results are specific to this 455,246-entry English list | Scope limit — no cross-linguistic or cross-corpus claim | The single claim this report stands behind is that the measurements in rows 3–5 are correct and reproducible from the deposited corpus. We do not claim they are new; if they reproduce a result already in the literature, we would simply have measured the same thing independently, and the earlier work has precedence. 9. Reproducibility and data availability The exact corpus is included as a data file with this Zenodo deposit. All tables are regenerated by a single script from that file (transition matrix, the N=2–16 atlas, Markov baseline, residuals, the interior-ID and shuffle controls, and the handover series). Random operations use a fixed seed. The consonant-side analysis of the same corpus is developed in the companion paper, "The Orthographic Junctions of English" (Zenodo 10.5281/zenodo.20430908); readers interested in the junction structure underlying these vowel profiles are referred there. --- Appendix A — Pointers for placing this in the literature Because we make no novelty claim, this appendix is not a gate to clear but a convenience for readers who want to situate these measurements. The queries below are starting points for finding related or overlapping work in the relevant databases; if any returns an equivalent result, that work simply predates ours. **Google Scholar** - `positional vowel distribution "word length" English lexicon` - `vowel proportion letter position orthographic corpus` - `"vowel mass" OR "vocalic proportion" position within word` - `letter position vowel consonant profile word length` - `"%V" orthographic text rhythm written` - `Markov chain vowel consonant English lexicon positional residual` - `consonant vowel skeleton position word length English` **ACL Anthology** (aclanthology.org search) - `positional informativeness word lexicon` - `word-initial information front-loading cross-linguistic` - `vowel consonant positional entropy lexicon` - `phonotactic position word length distribution` - Follow citation graph forward from Pimentel, Cotterell & Roark (2021), "Disambiguatory Signals are Stronger in Wordinitial Positions," **LingBuzz** (ling.auf.net) - `orthographic rhythm vowel positional` - `vowel distribution word length English` - `phonotactic edge effect lexicon position` - `consonant vowel alternation positional null model` **Targeted checks** - Is there a "developmental sequence by word length" treatment of positional letter statistics anywhere? (most positional work fixes a single length, e.g. 5-letter studies) - Has anyone used a V/C Markov chain specifically as a *null to subtract* for positional analysis (rather than as a generative or typological model)? - Has the front→back "handover" with word length been described under any name? These are pointers, not obligations. If a reader finds an equivalent treatment, the present report should be read as an independent reproduction of it, with precedence to the earlier work. --- References - Grabe, E., & Low, E. L. (2002). Durational variability in speech and the rhythm class hypothesis. *Laboratory Phonology* 7, 515–546. - Hafer, M. A., & Weiss, S. F. (1974). Word segmentation by letter successor varieties. *Information Storage and Retrieval* 10(11–12), 371–385. - Harris, Z. S. (1955). From phoneme to morpheme. *Language* 31(2), 190–222. - Hirst, D. (2009). The rhythm of text and the rhythm of utterances: from metrics to models. *Interspeech 2009*. - Markov, A. A. (1913/2006). An example of statistical investigation of the text *Eugene Onegin* concerning the connection of samples in chains. (English translation, *Science in Context* 19(4), 591–600.) - Pimentel, T., Cotterell, R., & Roark, B. (2021). Disambiguatory signals are stronger in word-initial positions. *EACL 2021*. (arXiv:2102.02183) - Ramus, F., Nespor, M., & Mehler, J. (1999). Correlates of linguistic rhythm in the speech signal. *Cognition* 73(3), 265–292. - Sabatini, A. M. (2026a). From graphemic dependence to lexical structure: a Markovian perspective on Dante's Commedia. arXiv:2604.22626 [cs.CL]. - Sabatini, A. M. (2026b). Markov reads Puškin, again: a statistical journey into the poetic world of Evgenij Onegin. arXiv:2604.20221 [cs.CL]. - Shannon, C. E. (1948). A mathematical theory of communication. *Bell System Technical Journal* 27. - Shannon, C. E. (1951). Prediction and entropy of printed English. *Bell System Technical Journal* 30(1), 50–64. - Boicho Dimitrov Temelakiev, Saxon Ventura Research Ltd (2026). The Orthographic Junctions of English (OJE). Zenodo. [companion paper] 10.5281/zenodo.20430908 and 10.5281/zenodo.20450297

提供机构:
Zenodo
创建时间:
2026-06-08
二维码
社区交流群
二维码
科研交流群
商业服务