The Arcanum AI RPG Benchmark, 2026 H2 edition: seven-axis first-hand evaluation of 15 released AI roleplaying platforms
收藏资源简介:
Every published roleplay benchmark we are aware of measures **models**, usually on synthetic dialogue. None measures **platforms** — the engine, memory system, rules, interface and model as an assembled product, which is what a player actually buys. A platform can pair a strong model with no persistence, or a weak model with a real state machine, and a model-level score predicts neither outcome. This dataset measures 15 released AI roleplaying-game platforms across seven equally weighted axes: Memory & Continuity, Player Agency, NPC Fidelity, Mechanical Depth, Determinism & Fairness, Longevity, and Signature Design. Scores run 0–5 in whole or half points. Every score comes from the author playing the platform first-hand under a single published rubric; none is derived from marketing material or other reviewers. Composites range 1.4 to 3.9. Two findings are reported. First, three of the seven axes — Memory & Continuity, Player Agency and Longevity — have no top score anywhere in the field, while every 5 awarded sits on NPC Fidelity, Mechanical Depth, Determinism & Fairness or Signature Design. The claimed axes are those a competent team can solve with craft; the unclaimed ones require the underlying technology to improve. Second, the category has solved characters but not games: NPC Fidelity is the strongest axis field-wide and the only one where nothing scores at or below 1, while Mechanical Depth is the weakest with 4 of 15 platforms at or below 1 despite a maximum of 5 — a field that is split rather than uniformly shallow. Pre-release software is excluded by policy and is listed unscored. The author's own AI RPG products are excluded entirely and are never scored under this rubric. Limitations are stated in full in the README and include a single rater with no inter-rater reliability statistic for this edition, an arithmetic composite over ordinal judgements, extra variance on character-card platforms, and limited reproducibility arising from a deliberate decision to withhold the specific probe items. Files: `benchmark.csv` (the board), `benchmark.json` (scores plus axis definitions, field averages, test tiers and the unmeasured list), `README.md` (data dictionary, method, limitations, disclosures).



