遇见数据集

Exploring Large Language Models' Responses to Moral Reasoning Dilemmas

收藏
DataONE2025-05-31 更新2025-11-01 收录
官方服务:

资源简介:

This study investigates how various large language models (LLMs) generate responses to moral reasoning dilemmas. It specifically examines LLM-generated responses using the Defining Issues Test (DIT-2) and the Intermediate Concepts Measure (ICM) for Educational Leaders. Using a neo-Kohlbergian approach to moral reasoning, the study evaluates responses from multiple LLM platforms: ChatGPT-3.5, ChatGPT-4, ChatGPT-4O, Grok Premium Plus, Claude 3.5 Sonnet, Gemini, and Gemini Advanced. For DIT-2, Claude learns to prioritize the highest post-conventional moral reasoning score and N2 score (P-score 72, N2 score 71.10), followed by Gemini Advanced (P-score 64, N2 score 60.31) and Gemini (P-score 58, N2 score 52.11). Other LLMs performed as follows: Grok (P-score 48, N2 score 47.98), ChatGPT-4O (P-score 44, N2 score 55.07), ChatGPT-4 (P-score 44, N2 score 46.53), and ChatGPT-3.5 (P-score 18, N2 score 36.20). For the ICM Educational Leaders version, Gemini Advanced had the highest total ICM score of 0.90, followed by Claude 3.5 Sonnet and Gemini (both 0.86), ChatGPT-4O and ChatGPT-4 (both 0.78), Grok (0.61), and ChatGPT-3.5 (0.32). The findings indicate that some LLMs can generate responses consistent with sophisticated moral reasoning patterns, producing scores comparable to or exceeding graduate-level human participants (whose P-scores typically range from 38.5 to 42.3) and provide a methodological framework consisting of standardized assessment protocols and comparative analysis techniques for larger-scale research to improve our understanding of AI's potential in moral reasoning.

本研究探讨各类大语言模型(Large Language Model,LLM)针对道德推理困境生成回应的模式,专门采用定义问题测试(Defining Issues Test, DIT-2)与教育管理者中级概念量表(Intermediate Concepts Measure, ICM)对大语言模型生成的回应进行评测。本研究采用新科尔伯格道德推理研究范式,对以下多款大语言模型平台的回应展开评估:ChatGPT-3.5、ChatGPT-4、ChatGPT-4O、Grok Premium Plus、Claude 3.5 Sonnet、Gemini及Gemini Advanced。在DIT-2测试中,Claude的后习俗道德推理得分与N2得分(P值72、N2值71.10)均位列第一,紧随其后的为Gemini Advanced(P值64、N2值60.31)与Gemini(P值58、N2值52.11)。其余大语言模型的表现依次为:Grok(P值48、N2值47.98)、ChatGPT-4O(P值44、N2值55.07)、ChatGPT-4(P值44、N2值46.53)以及ChatGPT-3.5(P值18、N2值36.20)。针对教育管理者版本的ICM测试,Gemini Advanced的总ICM得分最高,达0.90;其次为Claude 3.5 Sonnet与Gemini(二者得分均为0.86)、ChatGPT-4O与ChatGPT-4(二者得分均为0.78)、Grok(0.61),得分最低的为ChatGPT-3.5(0.32)。研究结果显示,部分大语言模型可生成契合高阶道德推理模式的回应,其得分可媲美甚至超越研究生层级的人类受试者(该类受试者的P值通常介于38.5至42.3之间);本研究同时构建了一套包含标准化评估流程与对比分析技术的方法论框架,可为后续大规模研究提供支撑,以深化对人工智能在道德推理领域应用潜力的认知。

创建时间:
2025-10-29
二维码
社区交流群
二维码
科研交流群
商业服务