Responses of Multimodal Large Language Models on BEMA, TUG-K, QMVI and FTGOT
收藏资源简介:
The dataset contains full responses of a selection of Multimodal Large Language Models on four physics concept inventories requiring the interpretation of images. The inventories are the Brief Electricity and Magnetism Assessment (BEMA), the Test of Understanding Graphs in Kinematics (TUG-K), the Quantum Mechanics Visual Instrument (QMVI) and the Four Tier Geometrical Optics Test (FTGOT). Each of the total 102 items was submitted 10 times. The models tested were: claude-opus-4-20250514, claude-sonnet-4-20250514, claude-3-5-haiku-20241022, gemini-2.5-pro-preview-06-05, gemini-2.5-flash-preview-05-20, gemini-2.0-flash, gemma-3-27b-it, gemma-3-4b-it, o3-2025-04-16, o4-mini-2025-04-16, gpt-4.1-2025-04-14, gpt-4.1-mini-2025-04-14, gpt-4.1-nano-2025-04-14, and gpt-4o-2024-11-20, gpt-5-2025-08-07, gpt-5-mini-2025-08-07, gpt-5-nano-2025-08-07.



