// benchmark
The Arcanum AI RPG Benchmark
2026 H2 edition · 16 platforms measured · last updated 2026-08-02 · doi:10.5281/zenodo.21721591
Across 16 platforms scored on seven axes, not one has reached the top of Memory & Continuity, Player Agency and Longevity. Every maximum score on this board belongs to NPC Fidelity, Mechanical Depth, Determinism & Fairness and Signature Design instead — and those are the craft problems, not the hard ones.
Every roleplay benchmark we can find measures models on synthetic dialogue. None measures platforms, which is what players actually buy — the engine, the memory system, the rules, and the model together. This page is our attempt at the missing one: 16 platforms, all played first-hand, all scored on the same published rubric.
It is not a shopping guide, and it deliberately doesn't rank anything for you. If that's what you came for, best AI RPG by priority names the winner on each axis, best AI roleplay platforms ranks by player type, and the compare table lets you sort the whole set yourself. This page asks a different question: after measuring the field, what is this category actually good and bad at?
What Is In It — and What Isn't
Everything measured here is a platform you can play today. Products still in closed alpha or closed beta are not scored — we don't rate pre-release software, because grading an unfinished build is unfair to it and an early verdict tends to outlive the version it described. So read this as a measurement of what's currently playable, not as the final state of the category.
Not yet measured (10): AIDungeonMaster.ai, Craft, Fable Forge, Myth Maker AI, MythEngyn, Questner, RoleForge, Taverna, Tidefall, Voyage. These are in closed alpha, closed beta, or application-gated early access. They carry no score and are listed alphabetically, not ranked — several are covered in our first looks, and each enters the board when it launches and earns a rating.
The rubric, the scale, the test protocol, and the disclosure rules are all on the methodology page. Scores run 0–5 per axis; the headline number is the mean of the axes that apply, never a holistic impression.
Finding 1 — Three Problems Nobody Has Solved
Three of the seven axes have no top score anywhere in the field: Memory & Continuity (best: 4), Player Agency (best: 4) and Longevity (best: 4). Every 5 we have awarded sits on NPC Fidelity, Mechanical Depth, Determinism & Fairness and Signature Design.
That split is not random. The axes where somebody has hit full marks are the ones a good team can solve with craft — write a real ruleset, adjudicate consistently, give the thing a voice, build characters that hold. The axes where nobody has are the ones that need the underlying technology to be better than it currently is: remembering what happened, letting a player act freely without the world falling apart, and staying coherent long enough for a campaign to matter.
If you have ever wondered why every AI RPG feels like it is fighting the same three battles, this is the measured version of that intuition. Why AI campaigns fall apart at turn 50 covers the mechanism behind the first one.
Finding 2 — The Category Has Solved Characters, Not Games
NPC Fidelity is the strongest axis field-wide, averaging 3.19 across 16 platforms. It is also the only axis where nothing scores at or below 1 — every platform we have tested can make a character feel like a person.
Mechanical Depth is the weakest, averaging 2.19, with 4 of 16 platforms scoring 1 or below — and yet its maximum is 5. That gap between the best and the average is the widest on the board, which tells you the field is not uniformly shallow so much as split: a small number of genuine game engines, and a long tail of chat windows with a fantasy skin.
Put together: this is a category that arrived from the chatbot side of the family and is still working its way toward being games. The parts it inherited are strong. The parts it has to build are not, yet.
Where the Field Stands, Axis by Axis
Mean is the field average across all 16 measured platforms. Best is the highest score any single platform has reached — where that is below 5, the axis is still unclaimed.
| Axis | The question | Mean | Best | Held by |
|---|---|---|---|---|
| NPC Fidelity | Consistent personality, own goals, durable reactions — or agreeable mirrors? | 3.19 | 5 | Janitor AI |
| Memory & Continuity | Does the world remember what happened, and for how long? | 2.66 | 4 · unclaimed | WyrdTale, Hidden Door |
| Signature Design | Does it do something nobody else does? | 2.63 | 5 | Friends & Fables |
| Determinism & Fairness | Outcomes earned and consistent, or arbitrary and luck-washed? | 2.56 | 5 | Friends & Fables |
| Longevity | Does it survive past turn 50 and the novelty window? | 2.31 | 4 · unclaimed | WyrdTale, Hidden Door |
| Player Agency | Railroading AND puppeting — can you act off-path, and does the narrator stay out of your character's mouth? | 2.25 | 4 · unclaimed | AI Dungeon |
| Mechanical Depth | Real systems underneath, or narration wearing dice? | 2.19 | 5 | Friends & Fables |
The Board
All 16 measured platforms, every axis, sorted by composite. The spread runs 1.4 to 3.9. A composite averages seven separate judgements and then tells you none of them, so read across the row rather than down the column — the axis winners are mostly not the platforms at the top.
| Platform | Score | Memory | Player | NPC | Mechanical | Determinism | Longevity | Signature |
|---|---|---|---|---|---|---|---|---|
| Friends & Fables | 3.9 | 3.5 | 3 | 4 | 5 | 5 | 2 | 5 |
| Janitor AI | 3.0 | 2 | 3 | 5 | 2 | 3 | 3 | 3 |
| Old Greg's Tavern | 3.0 | 3 | 3 | 3 | 2 | 3 | 3 | 4 |
| WyrdTale | 3.0 | 4 | 1 | 3 | 3 | 2 | 4 | 4 |
| MacerAI | 2.9 | 3 | 3 | 3 | 3 | 3 | 2 | 3 |
| Deep Realms | 2.7 | 3 | 3 | 3 | 2 | 3 | 3 | 2 |
| Hidden Door | 2.7 | 4 | 1 | 4 | 1 | 2 | 4 | 3 |
| Infinity DM | 2.7 | 3 | 3 | 3 | 2 | 3 | 2 | 3 |
| AI Dungeon | 2.6 | 3 | 4 | 3 | 2 | 2 | 2 | 2 |
| FableAI | 2.4 | 3 | 2 | 3 | 2 | 3 | 2 | 2 |
| AI Game Master | 2.3 | 3 | 2 | 3 | 2 | 2 | 1 | 3 |
| AI Realm | 2.3 | 1 | 2 | 3 | 4 | 3 | 1 | 2 |
| Character.AI | 2.1 | 1 | 2 | 4 | 1 | 2 | 2 | 3 |
| Chub AI | 1.9 | 2 | 2 | 2 | 1 | 2 | 3 | 1 |
| Dunia | 1.7 | 2 | 0 | 3 | 2 | 1 | 2 | 2 |
| Perchance | 1.4 | 2 | 2 | 2 | 1 | 2 | 1 | 0 |
Highlighted cells are the best score on that axis. Column headings are shortened — full axis names and the question each one answers are in the table above.
Take the Data
The whole board, in full: download the CSV or the JSON. Both are generated from the same source as the table above at build time, so the file and the page cannot drift apart. The JSON additionally carries the axis definitions, the field averages, the test tier each platform received, and the list of platforms not yet measured.
The dataset is also deposited in Zenodo, the open research repository run by CERN, which gives it a permanent DOI and an archive independent of this site. The deposit carries a full data dictionary, the method, and the limitations in longer form than this page has room for.
The archive is refreshed quarterly; this page is refreshed the moment a platform is scored. A platform enters the board as soon as it earns a rating, so between deposits the table above can be ahead of the archived snapshot — the current board was last updated 2026-08-02, and the deposited snapshot is dated 2026-07-29. That is deliberate. Cite the DOI when you need a number that cannot move under you; cite this page when you want the field as it stands today.
The data is published under CC BY 4.0. Reuse it, chart it, argue with it, reprint the entire table — commercially or otherwise. The one condition is attribution with a link back. We would rather be checked and named than kept tidy and ignored. (That licence covers this dataset. The articles elsewhere on the site stay under the usual terms.)
Cite it as:
The Arcanum AI RPG Benchmark, 2026 H2 edition (archived snapshot, 2026-07-29). Arcanum RPGs. Zenodo. https://doi.org/10.5281/zenodo.21721591 (CC BY 4.0).
That DOI pins the deposited 2026 H2 snapshot permanently, which is what you want when quoting a specific figure — a later deposit will move the numbers, and that is the point rather than a correction. If you mean the benchmark as an ongoing thing rather than this edition's numbers, cite 10.5281/zenodo.21721590 instead; it always resolves to the newest edition. Before republishing any of it, read the three limits below — they are the parts a table of numbers can't tell you on its own.
How to Read This Honestly
Three limits worth stating plainly, because a benchmark that won't name its own soft spots is asking for more trust than it has earned.
- It measures what shipped. The board is a snapshot of released software on the date above. It says nothing about what is in beta, and nothing about what a platform will be in six months.
- Card platforms carry extra variance. Where you supply a community-made card and the platform runs it, a share of the experience is the card rather than the product. We hold the card protocol constant so these entries stay comparable to each other, but their absolute numbers are softer than the rest. The full caveat is on the methodology page.
- One tester, one rubric. Every score comes from the same person playing the thing, which buys consistency and costs breadth. That is the honest trade of a first-hand benchmark over an aggregated one.
We publish the method and keep the probe items private, for the reason set out on the methodology page: a benchmark that publishes its test items stops measuring quality and starts measuring who read the test.
What Changed This Edition
This is the first edition, so there is nothing to compare it against yet. From the next one, this section carries the delta — which platforms entered, which scores moved and why, and whether any of the unclaimed axes finally got a top score.
There is, however, a before. Every platform on the board at launch carried a holistic rating first, set without a rubric, and rebuilding those on the seven axes moved every single one — mostly downward, and by as much as two points. Why our ratings changed publishes the full before-and-after, and why a lower number here almost never means a worse platform.
That is the part we expect to matter most over time. A single snapshot tells you where the field is; a run of them tells you whether it is actually getting better, which is a harder question and a more useful one.
Corrections
If a number here doesn't survive contact with your experience of the same product, we want to hear it — [email protected]. Disagreeing with a verdict isn't a correction, but a score whose reasoning doesn't hold up is.
Building one of these platforms? How we review, and what we won't sell.