WEP Elicitation Experiment: Unlikely ≈ Likely — A Negation Failure in LLM Interpretation of Verbal Probability
收藏资源简介:
Version 2 — updated 2026-07-24. Complete raw data, code, figures, cell-level summaries, and reproduction script for the revised manuscript "Unlikely ≈ Likely: A Negation Failure in Large Language Model Interpretation of Verbal Probability" (Discover Data, Submission ID 76f98a21). Five instruction-tuned model endpoints (llama-3.1-8b-instant, llama-3.3-70b-versatile, openai/gpt-oss-20b, openai/gpt-oss-120b, qwen/qwen3.6-27b) were evaluated on 16 verbal probability expressions across 120 scenarios in 12 domains under two elicitation frames. 11,840 planned elicitations completed (deduplicated). New in this version: cell-level summaries (320 rows for the main experiment, 48 rows for the temperature sweep), a self-contained reproduce_tables.py that reproduces Tables 7–13 from the raw logs, updated figures, manuscript build sources, and all raw responses (gzipped JSON Lines). Responses are keyed by (model, frame, temperature, repetition, item); earliest-successful deduplication identical to the manuscript.



