Gold set, evaluation harness and measurement logs for "Exploitability-Aware Web Intrusion Response with a Non-Authoritative LLM Advisory Channel"
收藏资源简介:
Supporting material for a multi-agent decision layer that converts web intrusion detections into cost- and exploitability-aware responses. The deposit contains the gold set of 63 cases spanning 435 events with the labelling rubric fixed before measurement; the evaluation harness that produces every reported figure, including the four ablation configurations; the suite of 63 scripted advisory proposals used to exercise the validator boundary; and the raw measurement logs from the containerised testbed. The gold set is synthetic: every case, including its source addresses, is a literal in build_goldset.py written to exercise a specific decision path. No case was captured from production traffic, and the deposit contains no personal data. Scope of the claims: the labelling rubric and the decision engine were designed by the same author, so the agreement figures measure implementation fidelity to that rubric rather than independent validation of the rubric itself.



