Top

The State of AI RPGs, September 2026: Solved at Turn One, Unsolved at Turn Fifty

We have scored 36 AI RPGs on the same seven axes, all from first-hand play, and the clearest thing in the data is not which product wins. It is that this category has solved every problem you can see in the first ten minutes and none of the problems you can only see in the tenth hour.

This is a dated edition. The figures below are what the field looked like on 15 September 2026, and we are going to leave them standing rather than quietly refresh them, because a state-of-the-field that silently updates is not a record of anything. When these numbers move, we will publish the next edition and you will be able to read them against each other. Live figures always sit on the benchmark, which recomputes itself.

A note on the numbers, added 18 September 2026. Everything below was true when this edition was published on 15 September, and it stays as written. One product has been scored since: Master of Dungeon, at 3.3, which takes the dataset to 37 products and the platform board to 25. It changes no ceiling on this page, because every one of its axes sits at 3 or 3.5, and it moves the 3.0–3.9 band from seven products to eight. Current figures are always on the benchmark.

Disclosures: several platforms in this dataset gave us free access or credits for review — Craft, Voyage, Tidefall, NOPOTIONS, Tabled, Tales RPG, aiga_, ArcQuill and AI Game Master. Fable Forge holds a paid sponsored placement on our directory and Questner bought an audit from Arcanum Forge, our design service; neither is scored here, because a rated platform may never be a client. None of it buys coverage or a score — how we handle this.

The Shape of the Field

Start with the distribution, because it is the part that recommendation lists structurally cannot show you. Every product below has been played and scored on seven axes; the headline figure is the mean of those axes.

Composite bandProductsShare
4.0 and above26%
3.0 – 3.9719%
2.0 – 2.92158%
Below 2.0617%

The field average is 2.61. Two products out of 36 clear 4. The large middle is not a middle in the sense of “decent” — it is a band of products that do one thing acceptably and three things poorly.

Then the figure that says the most about what it is like to browse this category: 18 of the 36 have no axis above 3. Exactly half the field is unremarkable at everything we know how to measure. The other half has at least one genuine strength, which is usually the thing its landing page is built around.

That is also the honest frame for the two boards. Hosted platforms average 2.83 across 24 products; prompt-games — a text file you paste into a subscription you already pay for — average 2.17 across 12. We keep them on separate boards for that reason. Averaging them produces a number that describes neither.

Finding One: Five Axes Have a Perfect Score. Two Never Have.

Here is the whole argument in one table. For each axis: the field average, the best score anyone has achieved, and how many products have reached full marks.

AxisField averageBest in fieldProducts at 5
Signature Design2.8655
Determinism & Fairness2.6054
NPC Fidelity2.9653
Mechanical Depth2.2452
Player Agency2.5051
Memory & Continuity2.694.250
Longevity2.4040

Two things are worth separating here, because they are usually confused.

An average tells you what you will probably get. A ceiling tells you what anybody has ever managed. Those are different questions and they give different answers. Mechanical Depth has the worst average in the field, 2.24, and yet two products have reached full marks on it — Craft by letting world authors write real systems in code, Friends & Fables by implementing an existing ruleset properly. Depth is not unsolved. It is a hard problem with two demonstrated solutions that almost nobody attempts, because attempting it means building an engine.

Memory is the opposite shape. Its average, 2.69, is the second highest of the seven — most products do something about memory, and the something is usually adequate for an evening. But in two years of testing, nothing has ever been awarded full marks. The ceiling is Voyage at Memory — 4.25, with Craft, Hidden Door, Tabled and WyrdTale behind it at Memory — 4. Memory is an easy problem to half-solve and one nobody has finished.

Longevity is the same, more starkly: an average of 2.40, a best in field of Longevity — 4 shared by Hidden Door, Voyage and WyrdTale, and no perfect score.

The clock behind the table

Now read that table’s right-hand column again, in order: 5, 4, 3, 2, 1, 0, 0.

We did not sort it that way to be cute; it fell out. And the order it fell into is, almost exactly, how long you have to play before the axis becomes visible.

  • Signature Design — the most-maxed axis. Visible in a screenshot.
  • Determinism & Fairness — visible within a few rolls.
  • NPC Fidelity — visible in the first real conversation.
  • Mechanical Depth — visible in the first hour, once you try to use a system.
  • Player Agency — visible the first time you deliberately go off-path.
  • Memory & Continuity — visible around turn fifty.
  • Longevity — visible only after the novelty has worn off, which is the point at which most people have stopped evaluating and started either playing or leaving.

The category’s competence has a clock on it. Everything a player can check before they subscribe has been solved by somebody; the two things they can only discover afterwards have been solved by nobody.

We want to be careful about how hard we lean on that, because these are small counts over 36 products and a single new entry could scuff the staircase. But the two ends of it are not a small-number artefact. Five separate products have reached the top of Signature Design. None has reached the top of Memory or Longevity, across every product we have ever scored.

Finding Two: The Field Solved Characters and Never Solved Consequences

NPC Fidelity is our strongest axis at 2.96. Mechanical Depth is our weakest at 2.24. That gap is the single most stable fact in this dataset, and there is a sharper version of it.

Four products score 4 or better on NPC Fidelity while scoring 2 or worse on Mechanical Depth: Character.AI, Hidden Door, Janitor AI and Tidefall. Those are products with characters who hold a personality, want things, and push back — running on top of almost nothing.

Not one product has ever done the reverse. There is no AI RPG in our dataset with real systems underneath and flat, agreeable characters.

That asymmetry is the tell. Good characters are close to free now — they are a gift of the underlying models, which is why four products can have them without building anything. Systems are not free, because a language model has nowhere to keep a number between turns except the conversation itself. 24 of the 36 score 2 or below on Mechanical Depth. Two thirds of this category is narration wearing dice.

The consequence for a player is specific, and it is not “the game is bad”. It is that the failure arrives late. A cast that behaves like people and a world that cannot remember what they did is a product that is wonderful on Friday and hollow by the following Thursday. Why AI campaigns fall apart at turn 50 covers the mechanism; why AI RPGs cannot do mechanics covers why the fix is architectural rather than a matter of writing a better prompt.

Finding Three: What Moved Since July

Our previous edition was published on 30 July 2026. Six weeks later, three things have changed at the top of the board, and almost nothing has changed in the middle.

  • The first composite above 4. Craft landed at 4.3 in late August, the first product to beat what had been the board’s ceiling. Voyage then passed it at 4.4 on 31 August.
  • The Memory ceiling was raised for the first time, from 4 to Memory — 4.25 (Voyage). It is a quarter point, which in our rubric is a tie-break and nothing else — it exists to separate Voyage from the products it was tied with, and the review names the finding that separates them.
  • The first perfect Player Agency score ever awarded. NOPOTIONS took Player Agency — 5 on 4 September, on an axis that had stood unclaimed since we started measuring.

What did not move: the field average, the shape of the distribution, and the two unclaimed axes. Improvement in this category is concentrated in a handful of products and is not diffusing outward. A player picking at random in September 2026 is having roughly the experience they would have had in July.

What We Got Wrong

The July edition of this dataset is deposited on Zenodo under a permanent DOI, and its README makes a falsifiable claim: that three of the seven axes — Memory, Longevity and Player Agency — had no top score anywhere in the field, and that the unclaimed ones were unclaimed because they needed the underlying technology to improve.

NOPOTIONS broke that on 4 September. And it broke the reasoning, not just the count. Player Agency did not fall to a better model. It fell to restraint — a narrator that describes a scene and then stops, on a map open enough that there is nowhere to railroad you back onto. That is craft, which means the axis was misfiled in July rather than merely unsolved.

We cannot edit the deposited file, which is the entire reason for depositing it. What a DOI is, and why our benchmark has one has the full account, including the sentence that turned out to be wrong.

It is worth stating what this implies for the two axes still standing. We said they were blocked on capability. We were wrong about that once already, in the same paragraph, about the axis nearest to them. Memory and Longevity may also turn out to be craft problems nobody has bothered with — and the fact that the products nearest the top on both got there by five different architectures, rather than by using a better model, is at least suggestive.

What Would Change Our Mind

Predictions, on the record, so the next edition can be checked against this one.

  1. Somebody will take Memory to full marks before Longevity. Memory has a measurable finish line — does the world know what happened, at turn 200 — and four products are within a point of it.
  2. The field average will not move more than 0.15. Improvement is concentrated at the top; the tail is made of products nobody is working on.
  3. The NPC-without-depth asymmetry will hold. If a product ever scores Mechanical Depth — 4 with NPC Fidelity — 2, the argument in Finding Two is wrong and we will say so here.

If any of those is wrong, it will be wrong in public, in the next edition, with this one still readable.

How We Measured This

The seven axes, the 0–5 scale, the quarter-point rule, the test tiers and the evidence standard are published in full on our methodology page. Every score comes from playing the product first-hand, by Rukka, against a fixed probe set — including deliberate attempts to do things a working system should refuse. We publish the method and not the probe content, because a published test measures who read the test.

Three limits worth stating plainly. Card-driven platforms carry extra variance, because part of what we are scoring is a community-written card rather than the product, and we publish that caveat rather than hiding it. We do not score pre-release software, so this measures what has shipped — several of the most interesting things in the category are in closed beta and carry no number here. And it is one tester on one rubric, which buys consistency and costs breadth.

One scoping note. The figures above cover all 36 scored products, platforms and prompt-games together, because this page is about the state of a category. The Arcanum AI RPG Benchmark publishes the platform half — 24 platforms, every axis — because it is a buyer’s instrument and a free prompt is not something you choose between. Per-axis figures there run slightly higher than the ones here. Where we have corrected a published figure, it is logged on corrections.

For the prior question — what separates a game from a very good chat, and why character quality is the worst guide to it — see what makes an AI RPG feel like a game.

If you want the seven problems explained one at a time rather than counted, AI RPG design: the seven problems is the companion to this page — and the axis deep-dives are mechanics, NPCs, agency, fairness and memory and longevity. If you want a product rather than an argument, best AI RPG by priority names the winner on each axis and the compare table lets you sort the field yourself.

Frequently Asked Questions

Are there any good AI RPGs?

Two of the 36 products we have scored average above 4 out of 5, and seven more sit between 3 and 3.9. So yes, but they are a small minority of what is on offer: 21 of the 36 land in the 2s, and 18 of them have no individual strength above 3 on any of our seven axes. The useful question is not whether a good one exists but which axis you care about, because the field is dramatically better at some of these problems than others.

What is the biggest weakness of AI RPGs in 2026?

Anything that only reveals itself late. Mechanical Depth is the worst axis by average at 2.24 out of 5, and Memory and Continuity and Longevity are the two axes where no product has ever been awarded full marks anywhere in the field. Those are the qualities you cannot judge from a trailer or a first session, which is why the category reviews far better than it plays.

Have AI RPGs actually improved?

Measurably, at the top, over roughly six weeks. Since our July edition the first composite above 4 was awarded, the Memory ceiling was raised from 4 to 4.25, and Player Agency received the first perfect score ever given on that axis. None of those moved the field average much, because improvement in this category is concentrated in a handful of products rather than spread across it.

Why are AI RPG characters so much better than AI RPG mechanics?

Because characters are what a language model is already good at and mechanics are what it is structurally bad at. NPC Fidelity is our strongest axis at 2.96 and Mechanical Depth our weakest at 2.24. Four products in our dataset score 4 or better on characters while scoring 2 or worse on systems, and not one has ever done the reverse. Good characters come close to free with the model; systems have to be built outside it.

How many AI RPGs are there?

More than anyone has counted, because the category has no agreed boundary and a new prompt-based one can exist in an afternoon. We track 32 platforms and 16 prompt-games, of which 36 have been played first-hand and scored. We do not claim that is the whole field, only that it is the part we have measured on one consistent rubric.

Do AI RPG ratings get worse over time?

Ours have, on average, and deliberately so. When we re-scored every platform against methodology v1.0 in July 2026, 13 of 15 went down, because a rubric that separates seven qualities is harsher than a single overall impression. A product that is a delight for an hour and thin for a campaign scores well on one axis and poorly on three.