WEP Elicitation Experiment: Verbal Probability Interpretation in Five LLMs
收藏资源简介:
Complete raw data, code, and stimuli for the manuscript "Unlikely ≈ Likely: A Negation Failure in Large Language Model Interpretation of Verbal Probability" (Discover Data, Submission ID 76f98a21-3619-46af-a486-4073b54c8eb0). Five instruction-tuned model endpoints spanning three families (llama-3.1-8b-instant, llama-3.3-70b-versatile, openai/gpt-oss-20b, openai/gpt-oss-120b, qwen/qwen3.6-27b) were evaluated on 16 verbal probability expressions across 120 scenarios in 12 domains under two elicitation frames. 11,840 planned elicitations; 11,829 completed (98.5% of cross-model panel, 100% of anchor model and temperature sweep). Includes: all raw API responses (including API failures and parse errors) in JSON Lines format, stimulus generation script with well-formedness checks, complete item and context files, analysis code, figure generation code, environment lock, SHA-256 checksums, and a 24-entry reference verification log. Responses are keyed by (model, frame, temperature, repetition, item); superseded records from interrupted runs are retained for audit.



