遇见数据集

Multi-Agent Behavioral Authentication Benchmark and Reproducibility Artifact

收藏
Zenodo2026-08-11 更新2026-08-13 收录
官方服务:

资源简介:

This version preserves the finalized synthetic benchmark corpus of 10,000 validated multi-agent LLM conversations and 19,216 five-turn authentication windows and the scientific evidence produced for the major revision of JSA-D-26-01060. The package includes the corrected 72-episode LangGraph framework-native pilot; deterministic and history-conditioned adaptive mimicry evaluations; target-aware MobileBERT, DeBERTa, frozen-probe, and SBERT audits; ONNX Runtime FP32 and dynamic INT8 MobileBERT artifacts; Raspberry Pi 5 gateway, service-memory, latency, concurrency, energy, and framework-event replay results; conversation-clustered statistical audits; and reproducibility scripts and notebooks. Version v1.1.3 separates scientific artifacts from journal-submission documents. It removes the clean and highlighted manuscripts, LaTeX submission source, point-by-point response, cover letter, highlights, and editor-specific review documents from the Zenodo package. Those files are supplied through Editorial Manager. Conversation text, labels, identities, splits, windows, predictions, metrics, model artifacts, experimental code behavior, and scientific conclusions are unchanged from v1.1.2. The package supports deterministic evaluation replay and artifact verification; it does not promise bit-identical regeneration of sampled or proprietary LLM outputs. Framework-native results show operating-threshold transfer failure and do not establish ranking retention. Adaptive comparisons use validation-selected matched false-accept-rate operating points. Raspberry Pi 5 results describe one board and software stack.

提供机构:
Zenodo
创建时间:
2026-08-11
二维码
社区交流群
二维码
科研交流群
商业服务