An in silico electorate: persona-based forecasting of the 2026 South Korean local elections, Pre-registered predictions and held-out evaluation data
收藏资源简介:
Pre-registered forecast. This deposit was time-stamped on 2026-06-03, prior to the publication of official results by the National Election Commission (NEC) for the 9th Korean nationwide local elections (제9회 전국동시지방선거). The prediction files are immutable and append-only: they constitute a falsifiable, pre-committed test of the method and will be scored against official returns. No file in this record may be edited post hoc; corrections, if any, are issued as new versioned deposits. Overview. We forecast the 2026 South Korean local elections with an in silico electorate: a population of synthetic voter personas whose political behaviour is simulated rather than polled. Each persona is mapped by a sovereign Korean large language model (KT Mi:dm 2.0, served locally) into a two-dimensional ideological space — economic/cultural conservatism on one axis, anti-establishment sentiment on the other — together with an abstention propensity. Personas cast simulated ballots by proximity to candidate positions (distance-weighted softmax with stochastic jitter and an abstention rule). District-level outcomes are then anchored to a climate target derived only from verified prior elections — the 2024 National Assembly election margin, blended with the 2025 presidential swing, plus a post-government-change honeymoon term (H = +10 pp toward the incoming party) — by solving for the electorate shift that realises that target. The persona simulation supplies the demographic composition of each result; the verified-election anchor supplies the climate. Crucial provenance distinction (read before scoring). The canonical, pre-registered model predictions use no opinion polls of any kind. Poll-informed files are included separately, clearly labelled, solely to enable a "did polls add value beyond the model" comparison — they are not the test of the model and must not be scored as such. File manifest results_local_ftt_pres.json — CANONICAL. Basic-municipality (기초단체장) forecast, 223 races. Model only, zero polls. Fields per race: predicted winner, party, D−R margin, raw persona margin, verified anchor components, electorate shift. results_metros_governors.json — CANONICAL. Metropolitan mayor / provincial governor (광역단체장) forecast, 16 races. Verified climate projection + documented candidate personal-vote residuals (Busan +7.4, Seoul −10.0). Zero polls. Internal held-out checks (provisional; official results supersede). Against the held-out poll set, the canonical no-poll model attains ~67% directional agreement at the basic-municipality level (mean absolute margin error ≈ 12 pp), against an estimated ceiling of ~72% for any province-level fundamentals model — the residual ~28% being candidate-level effects (incumbency, local notables) that fundamentals cannot observe. The provincial climate model matches the known pre-election state in 15 of 16 races. Coverage and limitations. Sejong and Jeju have no basic-municipality tier (single-tier administration) and are reported only at the metropolitan level; four newly created Incheon districts lack persona coverage and are omitted. The 14 National Assembly by-elections held concurrently are not included in this version. Predictions are layered (directional and margin) and never collapsed to a single index. Races off the Democratic–People-Power axis — chiefly in the Honam region, where the second pole is the Rebuilding Korea Party or independents — are flagged low-confidence, since the spatial model resolves only the two-major-party axis. Ethics and data. The persona dataset is fully synthetic: it contains no personally identifiable information and is never joined to real individuals. All inference runs on a sovereign Korean LLM, consistent with the project's data-sovereignty commitments. Confidence levels and low-signal districts are reported as first-class outputs rather than suppressed. Suggested keywords: computational social science; election forecasting; large language models; synthetic populations; agent-based modeling; South Korean politics; pre-registration.



