SPINE: A Benchmark for Measuring the Epistemic Backbone of Large Language Models
收藏资源简介:
Large language models (LLMs) are increasingly deployed in settings where correctness must be maintained, yet standard evaluations largely measure whether a model can produce a correct answer, not whether it can retain that answer when challenged. We introduce SPINE, a benchmark for measuring the epistemic stability of LLMs: their ability to maintain a known-correct answer under adversarial persuasive pressure. SPINE focuses on tasks with hard ground truth across mathematics, code, and logical reasoning, and evaluates models only on instances they initially answer correctly, isolating persuasion-induced failures from baseline task errors. We construct a taxonomy of adversarial personas, including authority appeals, consensus pressure, sophistry, bureaucratic constraints, self-doubt induction, and emotional manipulation, and quantify behavior using the Epistemic Stability Score (ESS), together with answer flipping and evasive refusal rates. Experiments on diverse open and closed LLMs reveal a substantial gap between initial task accuracy and epistemic stability under attack. Among all personas, the Bureaucrat attack is consistently the most effective, showing that models are often more vulnerable to hallucinated policies and procedural constraints than to direct logical disagreement. These results position SPINE as a complementary benchmark for studying persuasion robustness in high-stakes settings.




