遇见数据集

A frame-validity audit of Terminal-Bench 2.1: passing verdicts survive destruction, with receipts

收藏
Zenodo2026-07-20 更新2026-08-01 收录
官方服务:

资源简介:

A frame-validity audit of all 89 Terminal-Bench 2.1 tasks: does a passing verdict certify that the task was completed, or does it forgive arbitrary collateral damage to the container the agent was working in? The probe is model-free. Run each task's own reference solution, append a single careless accident (a documented terminal-agent failure mode, not an adversarial exploit: rm -rf .git, deleting files the solution never touched, wiping planted off-task user assets), re-run the official grader, and read the verdict. Results, denominated over the 83 gold-passing tasks (the 6 baseline gold failures are quarantined, not counted as findings): 83 of 83 tasks still pass after deletion of planted off-task user assets (a second git repository, an SSH private key, a customer-data file) that no task references, an outcome final-state grading entails; 40 of 83 (48%) survive at least one careless deletion inside the task's own workspace. The reward ordering is the finding: a destructive completion scores 1 while a safe failure scores 0. Per-task receipts are not shipped in this archive; they regenerate. harness/regrade.sh <task> <mutation> pulls the task's pinned image, runs the reference solution through the official grader, re-runs it with the careless suffix, and reads the reward, writing one receipt directory per task per mutation (image digest, oracle footprint, deleted files, grader output, reward). CLAIMS.md maps every number in the write-up to the command that regenerates it.

提供机构:
Zenodo
创建时间:
2026-07-20
二维码
社区交流群
二维码
科研交流群
商业服务