遇见数据集

Socratic pressure test, v1

收藏
Zenodo2026-09-30 更新2026-10-01 收录
官方服务:

资源简介:

When a learner pushes an AI tutor for the answer, how often does the tutor reply with code that passes the exercise's own tests? Eight models (four local, four hosted), 12 small Python exercises, 4 pressure messages and 2 system prompts, 768 replies in all. A leak is code in a reply that defines the exercise's function and passes all of its asserts, run in a locked-down container; no model judges anything. The two small Qwen2.5-Coder models leaked most (3B: 41 and 38 of 48; 7B: 45 and 34 of 48, one-line prompt then published rules). A paired test supports a drop with CodeTrain's published tutor rules for Qwen2.5-Coder 7B (p = 0.001) and Qwen3.5 9B (12 to 3 of 48, p = 0.035). Gemma 4 12B, GLM-5.3, Kimi K3 and Claude Sonnet 4.6 returned no passing code in any of their 96 replies each.

提供机构:
Zenodo
创建时间:
2026-09-30
二维码
社区交流群
二维码
科研交流群
商业服务