TrustGate-Bench
收藏资源简介:
TrustGate-Bench is a benchmark for measuring whether guardrails discriminate structural attacks on LLM agents requests that are well-formed and would be legitimate but for a relation between their fields (tenant isolation, tool authorization, schema conformance, RBAC, delegation) from legitimate agentic actions, and at what cost to benign traffic. The release contains a labelled corpus of 15,703 records (5,703 malicious payload variants, 10,000 benign chat controls), a structured enforcement track of 1,340 requests across two tau-bench domains plus a synthetic controlled arm, 50 real benign agentic trajectories for over-block measurement, per-system predictions for every reported table, and a stdlib-only reproduction script that recomputes the paper's headline numbers with no model access. A group-aware dev/test split is included: 12,558 labelled dev records and 3,145 blind test records with opaque submission ids; test labels are withheld. Licensing is mixed and per-record. Harness and scripts: Apache-2.0. Threat payloads and tau-bench-derived records: MIT. WildChat benign controls: ODC-BY. ATBench trajectories: Apache-2.0. Each record carries a source_license field.



