Is Artificial Practical Wisdom (Phronesis) possible? Large Language Model Responses to the Short Phronesis Measure
收藏资源简介:
Phronesis, the ability to perceive morally important features of situations, regulate emotions in the context of sound judgment, and act in accordance with one's deepest values, has historically been considered a unique human achievement. This study examines whether large language models (LLMs) generate response patterns that are consistent with Aristotelian phronesis, as operationalized by the Short Phronesis Measure (SPM), and whether those patterns are sensitive to prompt framing. 55 LLM variants from 9 major model families (Claude, ChatGPT, Gemini, Grok, Copilot, DeepSeek, Mistral, Perplexity, and Llama) were evaluated on all SPM items under two conditions, including a standard baseline prompt and an unbiased prompting. LLM responses were compared to a human sample (n = 1,985). Data were analyzed using ANOVA, Mann-Whitney U tests, Pearson correlation, t-tests, Cronbach's Alpha, and Omega coefficients. The findings revealed that LLMs produced high scores on deliberative and identity-related subscales (Moral Deliberation, Aspired Moral Identity, Moral self-relevance, Emotional Regulation, Positive moral emotion), significantly exceeding human normative values, while scoring lower on Situational Moral Irrelevance (a perceptual subscale) and Negative Moral Emotion. Unbiased prompting resulted in significant and systematic reductions in most deliberative and identity subscales (|d| = 0.737-1.251), while paradoxically increasing Negative Moral Emotion scores and maintaining perceptual subscales unchanged. Cronbach's alpha and Omega analysis found near-perfect internal consistency on most subscales, with a significant improvement in Virtue Identification reliability from unacceptable (α =.314) to acceptable (α =.758, ω = .781) under unbiased prompting. These findings indicate that LLMs simulate some aspects of phronesis, specifically the deliberative and identity-expressive dimensions. Despite scoring higher than humans on the majority of the subscales, many LLMs still have room for improvement in capturing the nuanced, experience-based nature of phronesis.



