Cross-tier Evaluation of LLMs for AI-Generated Teaching Exemplars: Dataset and Analysis Pipeline
收藏资源简介:
Reproducibility package for a study evaluating seven large language models (LLMs) on their ability to generate pedagogically valid teaching exemplars for professional clinical communication, against a same-rubric human-expert baseline. Includes raw rating data (467 records, 5 expert raters), 84 blinded LLM and 9 human-expert dialogue outputs, the full Python analysis pipeline, and all derived statistical outputs reported in the companion manuscript. Models evaluated: GPT-4o, GPT-4o-mini, Claude-3-Haiku, DeepSeek-v3, Qwen2.5-72B, Llama-3.3-70B, Qwen2.5-7B, plus 9 human-expert authored dialogues (3 anaesthesiologists × 3 segments). Domain: preoperative anaesthesia consultation as an extreme-case validation context. Five-dimension evaluation rubric grounded in cognitive apprenticeship, feedback literacy, and self-regulated learning theory. Companion manuscript submitted to Computers & Education, May 2026. Files are under embargo until manuscript acceptance to prevent pre-publication scoop; metadata and DOI are public for citation purposes.



