遇见数据集

Evaluating Large Language Models on the 2026 Korean CSAT Mathematics Exam: Measuring Mathematical Ability in a Zero–Data-Leakage Setting

收藏
Zenodo2026-02-09 更新2026-05-26 收录
官方服务:

资源简介:

This dataset contains inference outputs from evaluating 36+ state-of-the-art large language models (LLMs) on the 2026 Korean College Scholastic Ability Test (CSAT) Mathematics section (홀수형, Odd Type). The exam consists of 46 problems: 22 common problems (Math I & II) and 24 elective problems (8 each from Probability & Statistics, Calculus, and Geometry). The dataset includes 347 TSV files (3.72 MB total): 338 per-model result files, organized by experimental condition: Input modality: text_only, image_only, text_fig (text + figure) Prompt language: English (en), Korean (ko) Extended thinking: enabled (on), disabled (off) 9 ablation summary tables aggregating accuracy, score, latency, token usage, and cost across all conditions. Each result file records per-problem: model answer, correctness, response latency, prompt/completion token counts, and timestamps. Models evaluated include: OpenAI GPT-5 family (GPT-5, GPT-5-Mini, GPT-5-Nano, GPT-5-Codex), Anthropic Claude (Opus 4.1, Sonnet 4.5, Haiku 4.5), Google Gemini 2.5 (Pro, Flash), xAI Grok 4 (Grok 4, Grok 4-Fast), Qwen3-VL, Meta Llama 4 Maverick, DeepSeek R1/V3, Mistral Magistral, Kimi K2, MiniMax M2, Nvidia Nemotron, GLM-4.6, and others. Note on original exam data: The original CSAT exam problems (text, images, and answer keys) are not included in this deposit due to copyright restrictions by the Korea Institute for Curriculum and Evaluation (KICE). The exam materials can be obtained from the official KICE website (https://www.suneung.re.kr).

提供机构:
Zenodo
创建时间:
2026-02-09
二维码
社区交流群
二维码
科研交流群
商业服务