遇见数据集

Pact Benchmark: ICPC World Finals — Contract-First Multi-Agent vs Single-Agent Code Generation

收藏
Zenodo2026-03-04 更新2026-05-26 收录
官方服务:

资源简介:

Benchmark comparing Pact (contract-first multi-agent framework) against Claude Code on 5 ICPC World Finals competitive programming problems (212 test cases). Pact achieves 100% (212/212) vs Claude Code single-shot 79% (167/212) and iterative 92% (196/212). All conditions use Claude Opus 4.6. Includes test data, baseline results, full Pact state for both conditions (research and base), and reproduction scripts. The decisive problem is Trailing Digits (2020 World Finals): Claude Code scores 31/47 even with 5 retry iterations — the naive algorithm times out. Pact's interview and decomposition phases force upfront mathematical analysis, producing the correct O(log n) approach on the first attempt.

提供机构:
Zenodo
创建时间:
2026-03-04
二维码
社区交流群
二维码
科研交流群
商业服务