Allow, Hold, or Deny: Comparing the Speed, Accuracy, and Reported Certainty of Four AI Models — A comparative evaluation of CLM, JEV, KEV, and GLM 5.3 FlashX on agent-action authorization.
收藏资源简介:
A comparative evaluation of CLM, JEV, KEV, and GLM 5.3 FlashX on agent-action authorization. Version 1.0.0 of the reproducible evidence for a four-configuration comparison on 300 manually authored and reviewed agent-action authorization cases (89 allow, 111 hold, 100 deny). Fixed sequential order: CLM v0.1-8B with unmerged MPS implementation, JEV 1.13.0 hosted directly, KEV-4B on MLX, and GLM 5.3 FlashX through OpenRouter/Z.AI FP8. Each arm received five practices and 300 one-case scored requests. Correct decisions out of 300 were 111, 272, 273 and 260 respectively. CLM returned hold throughout. GLM produced 267 valid responses and 33 incomplete outputs at the frozen 640-token completion cap including reasoning. All failures are retained; no scored outputs were repaired, retried or relabeled. The archive includes preparation evidence, raw responses, cases and reference annotations, selected original frozen scientific snapshots, public approval provenance, normalized observations, workbook, figures, code, cost accounting and offline reproduction instructions. Comparison limits include the manually authored bank, one complete session per model, an unmerged CLM implementation with unverified CUDA parity, staged author-approved background-service waivers, an hours-long gap before GLM, differing local and hosted hardware/network/cache conditions, and native classifier versus self-reported probabilities. The evidence does not establish deployment safety or architecture causality. No Phase 2 or separate curiosity experiment is included. Original data, figures and documentation are CC BY 4.0. Original study code is MIT; third-party rights are retained. Private diagnostic and approval derivatives are explicitly marked with original and public hashes. No credentials, model weights or third-party documentation snapshots are redistributed.



