遇见数据集

Pisama-bench v1-lite: held-out evaluation set for multi-agent failure detectors

收藏
Zenodo2026-05-22 更新2026-05-26 收录
官方服务:

资源简介:

Pisama-bench v1-lite is the public held-out evaluation benchmark for the Pisama multi-agent failure detection system. The benchmark contains 1,774 entries across 66 detection types at hard difficulty, drawn deterministically from the internal Pisama golden dataset and capped at 30 entries per detection type for balanced per-detector evaluation. What this benchmark is for Outcome-only benchmarks (e.g. GAIA, SWE-Bench) tell you whether an agent finished its task. They do not tell you why a multi-agent system failed when it failed. Pisama-bench targets the failure-detection layer directly: given a trace fragment, can your detector flag the specific failure mode present? Per-detector F1 is the metric. Evaluation results (Pisama detectors, calibrated thresholds) Mean F1: 0.83 across 65 evaluated detectors Median F1: 0.85 20 of 65 detectors at F1 ≥ 0.90 44 of 65 detectors at F1 ≥ 0.80 55 of 65 detectors at F1 ≥ 0.70 10 detectors below F1 0.70: langgraph_state_corruption (0.36), langgraph_checkpoint_corruption (0.52), langgraph_tool_failure (0.55), decomposition (0.57), specification (0.57), delegation (0.62), langgraph_parallel_sync (0.62), completion (0.67), openclaw_tool_abuse (0.67), persona_drift (0.69). Disclosed honestly as recall gaps on hard cases. Cost: 0 LLM tokens to run the heuristic tier Reproducibility Evaluation script and per-detector results JSON ship in the benchmark directory. Anyone can rerun: cd backend && ./.venv/bin/python scripts/evaluate_v1_lite.py Comparison to related public benchmarks TRAIL (Patronus, 2025): 148 traces, 841 errors, 20+ failure types. Closest comparison. v1-lite is roughly 12× larger and explicitly per-detector stratified. Who&When (Ye et al., ICML 2025): 58 hand-crafted cases for agent-level attribution. Different task. Schema Each entry: id, detection_type, input_data (detector-specific shape), expected_detected (boolean), difficulty ("hard" for all v1-lite entries), source ("llm_generated" or "manual"), tags, created_at. Full schema documented in SCHEMA.md. License The benchmark payload (pisama-bench-v1-lite.json) is licensed under Pisama-Benchmark-1.0: derived works permitted for research and evaluation use; attribution required. See LICENSE.md for the full text. CC-BY-4.0 will be adopted once per-entry source attribution lands. Related work Companion methodology paper: Tiered Detection of Multi-Agent LLM Failures: An Empirical Calibration on TRAIL and Who&When (Nikulainen 2026, DOI 10.5281/zenodo.20091432). Code: github.com/tn-pisama/pisama. Citation @misc{pisama-bench-v1-lite, author = {Nikulainen, Tuomo}, title = {Pisama-bench v1-lite: held-out evaluation set for multi-agent failure detectors}, year = {2026}, publisher = {Zenodo}, doi = {[assigned on publish]} }

提供机构:
Zenodo
创建时间:
2026-05-22
二维码
社区交流群
二维码
科研交流群
商业服务