AI RPG Design: The Seven Problems, Ranked by How Badly the Field Handles Them
Designing an AI RPG is not one problem. It is seven, they are genuinely separable, and the field is dramatically better at some than others.
We know because we measured it. Every product in our directory is scored on seven axes from first-hand play — 28 products now, platforms and prompt-games alike, all on the same rubric. Lining up the field averages gives something more useful than a ranking of products: a ranking of the problems themselves, in order of how badly the category currently handles each one.
| # | Design problem | Field average | State of the art |
|---|---|---|---|
| 1 | Mechanical Depth | 2.00 | Two products above 3 |
| 2 | Longevity | 2.18 | Nothing above 4 |
| 3 | Player Agency | 2.25 | Nothing above 4 |
| 4 | Determinism & Fairness | 2.39 | One product above 3 |
| 5 | Memory & Continuity | 2.48 | Nothing above 4 |
| 6 | Signature Design | 2.50 | Rare but achievable |
| 7 | NPC Fidelity | 2.86 | Effectively solved |
This page is about how to think about each one. It is not a ranking of products — the benchmark reports on the field, the compare table sorts it by any axis, and best AI roleplay platforms recommends by player type. And if you want the beginner’s version of what an AI dungeon master even is, start here instead.
1. Mechanical Depth — the hardest problem in the category
The question: are there real systems underneath, or is this narration wearing dice?
The worst-performing axis by a clear margin, and the one where the gap between marketing and reality is widest. Eight of 28 products have effectively no mechanical layer at all. Only two score above 3.
The reason is not developer laziness. A language model has nowhere to put a number between turns except the conversation, so the value persists exactly as long as the text mentioning it stays within reach of attention. Plausibility beats arithmetic: if the scene has become desperate, a desperate amount of remaining health fits, and fitting is what the model optimises.
The exceptions prove the shape of the solution. The two products above 3 didn’t write better prompts — they moved resolution outside the model. Full breakdown: why AI RPGs can’t do mechanics.
2. Longevity — the one everybody underestimates
The question: does it survive past turn 50 and the novelty window?
Second-worst, and the axis most likely to be discovered after launch. Nearly everything in this category is good for an hour. The interesting question is what it is like in session six, and the honest answer for most of the field is: noticeably worse, in ways the first session gave no warning about.
There is a rule we apply here that binds our own scoring: when the engine is absent and the model alone carries the story and the memory, longevity cannot exceed 3, however good the first hours are. Models drift across a long campaign in ways only a system can catch. It’s also why ambition can backfire — the more a design asks a prompt to hold, the sooner it drops something.
Where this is ranked and measured: best AI RPG for a long campaign.
3. Player Agency — two failures wearing one name
The question: can you act off-path, and does the narrator stay out of your character’s mouth?
The most compressed of the seven: 15 of 28 products score exactly 2, and nothing has reached 5.
The design trap here is assuming this is one problem. Railroading — the world steering you back onto the intended path — is the one everybody scores. Puppeting — the narrator acting or speaking as your character — is the one almost nobody does, and it is the more common failure. It hides inside generous games, because permissiveness and authorship are different properties. A product can impose no restrictions whatsoever and still take the pen out of your hand for a sentence at a time.
Full breakdown: railroading and puppeting.
4. Determinism & Fairness — the axis with a hole in it
The question: are outcomes earned and consistent, or arbitrary and luck-washed?
The distribution here is the strangest in the dataset: one product scores 5, nothing scores 4, and then there’s a cliff. Twenty-five of 28 sit at 2 or 3.
Models default to yes. Refusal is rarer in training data, harder to write well, and reads as unhelpful — and in most products there is no state underneath from which a refusal could be justified anyway. The result is a game that agrees with you, which feels like success for about an hour.
The tempting fix is to restrict the player, and our data says it doesn’t work: agency and fairness correlate positively across the field, and the two most restrictive products are among the least fair. Full breakdown: your AI dungeon master is too generous.
5. Memory & Continuity — the famous one
The question: does the world remember what happened, and for how long?
The problem everyone in this category already knows about, which is why it isn’t first on this list. Nothing has scored 5, and a 5 would mean a campaign that holds together indefinitely without you doing maintenance. Every product here is managing a hard limit rather than removing it. The good implementations delay the failure; none prevent it.
The mechanism, and the fixes that actually work: why your AI campaign falls apart at turn 50.
6. Signature Design — the cheapest axis to win
The question: does it do something nobody else does?
Worth calling out because it is the one axis where a small team is not structurally disadvantaged. It requires no infrastructure — only a decision. The highest signature scores in our games directory belong to a dynasty game where your ruler ages and dies every five-year turn, and an isekai world attempting base building and beast taming at once. Neither has notable engineering behind it. Both are doing something the rest of the field isn’t.
The trap is that signature ideas often cost you on other axes: the five-year turn that makes that dynasty game distinctive is also precisely why it forgets its own history. A distinctive idea is not free, and it is worth knowing which axis it is charging.
7. NPC Fidelity — effectively solved, and mostly not by you
The question: do characters hold a personality, pursue their own goals, and react durably?
The field’s strongest axis, and the only one where almost nothing is broken. But this is largely a gift of the underlying models rather than a differentiator anyone built — writing a convincing person is what these systems are best at.
Two things are still genuinely open. Casts flatten: put five characters in a scene and they converge into one narrator wearing several names. And almost nobody has NPCs with initiative — characters who want something you haven’t offered and pursue it while you’re elsewhere. That second gap is why dynamic factions barely exist in this category.
Full breakdown: why characters hold one-on-one and collapse in a party.
The Through-Line: Three Problems, One Missing Piece
Write all four of those breakdowns and something becomes obvious that no single one of them shows.
Mechanics need somewhere to keep a number between turns. NPCs with independent goals need somewhere to keep an agenda that advances while you’re elsewhere. Quests that survive a player going off-path need to be defined by state rather than steps, which means keeping the state.
Three of the seven problems are the same problem wearing different clothes: somewhere to keep things that isn’t the conversation.
That reframes the build question usefully. The instinct when an AI RPG feels thin is to write a better prompt, and prompts do help — they are the cheapest intervention available and they measurably slow every one of these failures. What they cannot do is create a record. Everything a prompt establishes lives in the same finite, shifting window as everything else, so it degrades on the same curve as the plot it was meant to govern.
You can build a genuinely enjoyable AI RPG without solving this. You cannot build one that remembers what it decided an hour ago — and the axis scores show exactly where that ceiling sits.
How We Measured This
The seven axes, the 0–5 scale, the test tiers and the evidence standard are published in full on our methodology page. Every score comes from playing the product, first-hand, by Rukka, against a fixed probe set — including deliberate attempts to do things a working system should refuse.
Three limits worth stating. Card-driven platforms carry extra variance, because part of what we score is the community card rather than the product, and we publish that caveat rather than hiding it. We don’t score pre-release software, so this measures what shipped. And it is one tester on one rubric, which buys consistency and costs breadth.
One scoping note on the numbers above: they cover all 28 scored products, platforms and prompt-games together, because this page is about what is hard to build. The Arcanum AI RPG Benchmark publishes the platform half of that — the 16 scored platforms, every axis — because it is a buyer’s instrument, and a free prompt isn’t something you choose between. The per-axis figures there are consequently a little higher than the ones on this page.
Building one of these? How we review, and what we won’t sell.
Frequently Asked Questions
What makes a good AI RPG? Seven things, and they are genuinely separable: whether the world remembers, whether you can act freely and stay the author of your character, whether characters hold their own personalities and goals, whether real systems run underneath, whether outcomes are earned rather than arbitrary, whether it survives past the novelty window, and whether it does anything nobody else does. Products are usually good at some and poor at others, which is why a single overall rating tells you very little about whether one suits you.
What is the hardest part of building an AI RPG? Mechanical depth, by our measurements. It averages 2.00 out of 5 across 28 products, the lowest of the seven design problems, and eight of them have effectively no systems underneath the story at all. The reason is architectural rather than a matter of effort: a language model has nowhere to keep a number between turns except the conversation itself, so a rules layer built purely in the prompt decays at the same rate as everything else in the context window.
Why do most AI RPGs feel similar? Because most of them are solving the same one problem — generating good prose in response to what you type — and leaving the other six to the model. Character writing is the field’s strongest area and it is largely a gift of the underlying models rather than a differentiator anyone built. The products that feel genuinely different are the ones that added something the model cannot do on its own, which almost always means keeping state outside the conversation.
Do you need an engine to build an AI RPG? Not to build one people enjoy, but yes to score well on several of the axes. Three of the seven problems — mechanics, NPCs with independent goals, and quests that survive a player going off-path — all bottleneck on the same requirement, which is somewhere to keep state that isn’t the conversation. You can build a genuinely good experience without it. You cannot build one that remembers what it decided an hour ago.
What is the difference between player freedom and player agency? Freedom is how much the game permits. Agency is whether you remain the author of your own character. They come apart more often than people expect: a game can impose no restrictions at all and still narrate your character doing things you never typed, which takes authorship away just as effectively as blocking the route. We score both failures on one axis for that reason.