Multi-Agent Behavioral Authentication Benchmark and Reproducibility Artifact
收藏资源简介:
This version preserves the finalized synthetic benchmark corpus of 10,000 validated multi-agent LLM conversations and 19,216 five-turn authentication windows and adds the evidence produced for the major revision of manuscript JSA-D-26-01060. New material includes a counterbalanced 72-episode LangGraph framework-native pilot; deterministic and history-conditioned adaptive mimicry evaluations; target-aware MobileBERT, DeBERTa, frozen-probe, and SBERT audits; ONNX Runtime FP32 and dynamic INT8 MobileBERT artifacts; a long-running Raspberry Pi 5 gateway with service-memory, latency, concurrency, energy, and framework-event replay results; the revised manuscript and point-by-point response; and a deterministic conversation-clustered statistical sensitivity audit. The package supports deterministic evaluation replay and artifact verification. It does not promise bit-identical regeneration of conversations or LLM rewrites because GPT-4.1 and Qwen-based generation can vary across model and runtime versions. Archived conversations, rewrites, probabilities, predictions, metrics, configurations, model archives, and hardware traces are the citable experimental objects. The revision evidence has explicit scope limits. The framework pilot exhibits operating-threshold transfer failure and does not establish ranking retention. Adaptive comparisons use validation-selected matched false-accept-rate operating points and do not establish a universal robustness ranking. Raspberry Pi 5 results describe one board and software stack.



