遇见数据集

Source code.

收藏
Figshare2026-02-27 更新2026-04-28 收录
官方服务:

资源简介:

IntroductionWhile LLMs are used to generate medical and dental MCQs, their alignment with Bloom’s Taxonomy remains unexplored.Materials and MethodsFive widely used LLMs, including ChatGPT-4o (OpenAI), Copilot Pro (Microsoft), Claude Sonnet 4 (Anthropic), Grok 3 (xAI), and DeepSeek R1 (DeepSeek) were evaluated. Each model generated 60 MCQs (total 300) based on content from an oral and maxillofacial anatomy textbook across the five cognitive levels of Bloom’s Taxonomy. Two independent investigators assessed each item using a 5-point Likert scale for remembering, understanding, applying, analyzing, and evaluating/creating. Inter-rater reliability was measured using weighted Cohen’s kappa. Model performance and inter-model differences were analyzed using the Kruskal–Wallis test.ResultsInter-rater reliability was moderate to strong (kappa = 0.74–0.86). Median scores for remembering, understanding, applying, and evaluating/creating were above 4 across all LLMs, while the analyzing level scored a median of 3.5 for ChatGPT-4o and DeepSeek R1. No significant difference was found between models in remembering and understanding levels (p > 0.05). Claude Sonnet 4 outperformed the other models at the applying, analyzing, and evaluating/creating levels (p = 0.01, 0.003, and 0.005, respectively). Within-model analysis showed that only Copilot Pro and Claude Sonnet 4 consistently aligned with Bloom’s cognitive levels across all categories. In contrast, ChatGPT-4o, DeepSeek R1, and Grok 3 performed significantly better at the lower cognitive levels (p = 0.00, 0.00, and 0.001, respectively).ConclusionsAll LLMs performed well at lower cognitive levels, while Claude Sonnet 4 achieved the highest alignment at higher-order levels.

引言 尽管大语言模型(Large Language Models,LLMs)已被用于生成医学与牙科多项选择题(Multiple Choice Questions,MCQs),但其与布鲁姆认知层级分类(Bloom’s Taxonomy)的契合程度仍有待深入探究。 材料与方法 本研究评估了5款广泛应用的大语言模型,包括ChatGPT-4o(OpenAI)、Copilot Pro(Microsoft)、Claude Sonnet 4(Anthropic)、Grok 3(xAI)以及DeepSeek R1(DeepSeek)。所有模型均基于口腔颌面解剖学教科书的内容,针对布鲁姆认知层级分类的5个认知维度,各生成60道多项选择题,总计生成300道题目。由两名独立研究者采用5级李克特量表(Likert scale),针对记忆、理解、应用、分析与评价/创造5个维度对每道题目进行评分。采用加权科恩kappa系数(weighted Cohen’s kappa)评估评分者间信度。使用克鲁斯卡尔-沃利斯检验(Kruskal–Wallis test)分析模型性能及模型间的性能差异。 结果 评分者间信度处于中等至较强水平(kappa值为0.74~0.86)。所有大语言模型在记忆、理解、应用及评价/创造维度的得分中位数均高于4分;而在分析维度,ChatGPT-4o与DeepSeek R1的得分中位数为3.5分。在记忆与理解维度,各模型间未发现显著性能差异(p > 0.05)。在应用、分析及评价/创造维度,Claude Sonnet 4的表现均优于其余模型(对应p值分别为0.01、0.003及0.005)。模型内分析显示,仅Copilot Pro与Claude Sonnet 4在所有认知类别中均与布鲁姆认知层级保持一致。与之相对,ChatGPT-4o、DeepSeek R1与Grok 3在低阶认知维度上的表现显著更优(对应p值分别为0.00、0.00及0.001)。 结论 所有大语言模型在低阶认知维度上均表现良好,而Claude Sonnet 4在高阶认知维度上的契合度最高。

创建时间:
2026-02-27
二维码
社区交流群
二维码
科研交流群
商业服务