WisdomBench-V5: Token-Balanced Matched-Failure Benchmark and Reviewer Artifact
收藏资源简介:
Reviewer artifact and benchmark documentation for a token-balanced matched-failure evaluation of language-agent context-state adaptation. The release contains reproducibility code, frozen configurations, sanitized score rows, request and response provenance, retry and replacement ledgers, sensitivity analyses, a data card, licensing inventory, and dataset documentation. The confirmatory execution contains 59 completed shards, 11,328 score rows, 480 prospectively generated held-out episodes, and two official online model routes. One GLM shard terminated by provider length limits; 48 omitted opportunities are explicitly retained in the predeclared missing-not-at-random analysis. The release does not claim base-model weight learning, human-like wisdom, deployment safety, or universal memory superiority.



