Cross-verifier failure-mode corpus for AI-agent evidence packages
收藏资源简介:
Version 0.2 (25 July 2026) preserves the complete v0.1 corpus and adds the processed 13-scenario carrier tables used by the related XATF and RAISE analyses. carrier/scenario-failure-matrix.csv records predeclared and observed carrier/agent verdicts and typed boundary attribution. carrier/scenario-metrics.csv records synthetic sizes, timings, seven audit-field indicators, and their 0--7 sum. scripts/verify-carrier-tables.py checks all 13 identifiers, expected-versus-observed verdicts, result labels, relative provenance paths, and audit-score arithmetic. README-v0.2.md and PROVENANCE-v0.2.md state the source hashes, transformation, related-work boundaries, and the evidentiary limits of the processed tables. The archive is deterministic and has SHA-256 27bc494e12f5f5964e21e21eba8ed5403875bbed3129abf492cf0c4ec24956e4. It contains processed synthetic results, not live X-Road traffic, credentials, private keys, or source-run logs. The carrier scenario set and carrier-binding profile belong to the related XATF work; the related RAISE paper uses the frozen tables for a distinct predeclared-oracle analysis. Supporting data for the article "Freshness Is Not Authenticity: Cross-Verifier Failure Modes and a Layered Verdict Model for AI-Agent Evidence Packages" (Anton Sokolov, Tyche Institute). The article uses an open agent-attestation reference implementation (EATF / Agent Evidence Package) as the system under test. This is the sanitised evidence behind the quantitative results in Section 5: baseline conformance vectors; three full-stack export runs exposing a canonical-byte contract disagreement (a fresh, internally valid package fails independent offline verification while a timestamped ledger reports valid); a 29-case / 87-observation edge-forge mutation corpus with cross-verifier disagreement; an agent-card binding wave; a carrier-boundary bench; and a 100-row format census. Private hostnames, IP addresses, absolute local paths, and development credentials have been redacted; processed result tables, case definitions, checksums, and runner scripts are retained so the reported verdicts can be re-derived. See README.md for the file-to-table mapping.



