Axiom 2026 evaluation dataset
收藏资源简介:
Evaluation records and reproduction scripts for T. Chaudhry, "From Context to Code: Codifying Cross-System Reasoning for LLM-Assisted Operations," IEEE Access, 2026. Contains the 270-cell dataset (9 models × 3 passes × 5 queries × 2 conditions) with primary-judge verdicts, correctness rubrics, second-rater (DeepSeek) verdicts for every cell, per-run harness configuration, per-cell tool-usage flags, the baseline-condition skill documents, and scripts that regenerate every quantitative result in the paper. v1.1.0 adds the second-rater verdicts and agreement script (paper §VI-B), harness configuration, tool-usage flags, and baseline skills; no cell was re-run and no primary verdict changed. Raw agent traces are withheld because they contain customer identifiers from a production system.



