Multi-Agent Behavioral Authentication Benchmark and Reproducibility Artifact
收藏资源简介:
This version preserves the finalized synthetic benchmark corpus of 10,000 validated multi-agent LLM conversations and 19,216 five-turn authentication windows and the evidence produced for the major revision of manuscript JSA-D-26-01060. The package includes the corrected 72-episode LangGraph framework-native pilot; deterministic and history-conditioned adaptive mimicry evaluations; target-aware MobileBERT, DeBERTa, frozen-probe, and SBERT audits; ONNX Runtime FP32 and dynamic INT8 MobileBERT artifacts; Raspberry Pi 5 gateway, service-memory, latency, concurrency, energy, and framework-event replay results; the revised manuscript; and the point-by-point response. Version 1.1.1 is a provenance-normalization patch. Historical implementation-route labels attached to deterministic template repairs and portable bundle copies are replaced by accurate release-safe metadata. Conversation text, labels, identities, splits, windows, predictions, metrics, model artifacts, and reviewer-facing scientific conclusions are unchanged from version 1.1.0. The package supports deterministic evaluation replay and artifact verification. It does not promise bit-identical regeneration of sampled or proprietary LLM outputs. The framework pilot exhibits operating-threshold transfer failure and does not establish ranking retention. Adaptive comparisons use validation-selected matched false-accept-rate operating points. Raspberry Pi 5 results describe one board and software stack.



