What Makes an AI RPG Actually Feel Like a Game? The Line Between a Story and a System
You can spend three hours talking to an AI inside a fantasy world and have a genuinely wonderful time. The uncomfortable question is whether you were playing a game.
A language model can already produce characters, quests, combat, lore, dialogue, surprises and entire adventures. None of that automatically makes a game. A chatbot can play a dungeon master extremely convincingly while there is essentially no game underneath it — and the better the writing gets, the harder that is to notice from inside a session.
This is our attempt to draw the line properly, because the category is still defined mostly by vibes. We have scored 37 products on seven axes from first-hand play, which gives us something better than an opinion about where the line sits.
Disclosures: several platforms named here gave us free access or credits for review — Craft, Voyage, Tidefall, NOPOTIONS, Tabled, Tales RPG, aiga_, ArcQuill and AI Game Master. None of it buys coverage or a score — how we handle this.
Narration Versus Simulation
Here is the whole distinction in one exchange.
You type: “I attack the dragon.”
The AI replies: “Your sword strikes the dragon’s scales. It roars in fury, and the cavern shakes.”
It sounds like a game. But what actually happened?
Possibly nothing. There may have been no attack roll, no hit points, no position, no armour value, no injury that persists past the paragraph, no possibility of failure, and no consequence that outlives the sentence. The model produced the outcome you implied, dressed well.
Now the same input in a system. It resolves your weapon against the dragon’s defences, checks your skill, your distance, your current hit points, a probability, a damage figure — and then the world changes to match. The dragon is wounded, or you are, and both facts survive the scene.
Narration describes what happens. Systems determine what can happen.
That is the line, and everything below is a way of testing which side of it you are standing on.
One: It Has to Be Able to Say No
This is the single most important property, and the easiest to test.
A game needs to be able to tell you no. You try to pick the king’s pocket and you fail. You attack the guard and you die. You steal the horse and the owner notices. You kill an NPC and they are actually gone. You spend your money and you cannot spend it again.
The test: can the system meaningfully stop you from getting what you want?
If it cannot, you are in interactive fiction with a very good cast. That is a real and enjoyable form — it is simply not an RPG, and the distinction matters when you are deciding what to pay for.
We measure this as Determinism & Fairness, and the field’s distribution is stark. Across 37 scored products:
| Fairness score | Products |
|---|---|
| 1 | 2 |
| 2 | 19 |
| 3 | 11 |
| 3.5 | 1 |
| 4 | 0 |
| 5 | 4 |
Nineteen of thirty-seven sit at exactly 2. There is a hole at 4 — nothing occupies the step between “inconsistent” and “the outcomes are genuinely earned.” Why models default to agreeing with you, and the four tells that give it away in play, are in your AI dungeon master is too generous.
Two: Persistent State, Not Just Memory
A game does not remember that “earlier in the story, the character acquired some money.” A game knows you have 37 gold.
Those are radically different things, and the difference is the reason so many AI campaigns feel like they are dissolving. There is a hierarchy here worth naming:
- Memory — “the king hates me.” The model recalls a fact about the fiction.
- Persistent state — “the king has banned me from the capital.” The fact has become a value the system holds.
- Simulation — the guards at the gate actually turn you away. The value is being read by something that acts on it.
The further down you go, the more like a game it becomes. Most AI RPGs live at the top of that list, which is why 15 of the 37 score 2 or below on Memory & Continuity, and why nothing in the field has ever been awarded full marks on that axis — the ceiling is 4.25. Why AI campaigns fall apart at turn 50 covers the mechanism underneath.
The test: does the world remember facts, or does it hold state?
Three: Rules That Exist Outside the Prose
If the model can change the rules whenever you ask, the rules were never rules.
Watch it happen:
“I jump across the canyon.” “Against all odds, you make it.”
“I jump across the five-hundred-metre canyon.” “You summon incredible strength and clear the leap.”
The system is protecting the narrative. It has decided that the satisfying paragraph matters more than the consistent world — which is exactly what a language model optimising for the next token will do unless something stops it.
A game has friction. Your character has limits, the world has limits, and neither bends because you would prefer a cooler story.
A game needs an authority greater than the narrator. In an AI RPG that authority has to live alongside the model rather than inside its improvisation, because a rule that exists only in the prompt degrades on the same curve as everything else in the context window. This is the field’s hardest problem and its worst axis: Mechanical Depth averages 2.27 out of 5, and 24 of the 37 score 2 or below. Two thirds of this category is narration wearing dice. Why AI RPGs cannot do mechanics is the full argument.
Four: Agency Is Not Choice
An AI offers you the tavern, the blacksmith, or the castle. That is choice.
Agency is burning the tavern down. It is refusing the quest and having the quest giver’s behaviour actually change. It is framing the merchant for the murder and watching an investigation form around what you did.
The test: can you do something nobody anticipated, and does the world deal with it coherently?
This is the one place where AI RPGs hold a structural advantage over traditional games. A hand-authored RPG can only respond to what a designer thought of in advance; a language model can improvise a response to literally anything. That advantage is real — and it only counts when there is a system underneath able to record what you did, or the improvisation evaporates by the next scene.
It is also the axis where the category has most recently proved something is possible. NOPOTIONS holds Player Agency — 5, the only perfect score ever awarded on that axis, and it got there by restraint rather than by a better model: a narrator that describes the scene and then stops, on a map open enough that there is nowhere to railroad you back onto. The two ways a game takes your character away from you — railroading and puppeting — are separated in railroading and puppeting.
Five: Does the World Exist When You Are Not Looking?
Ask an AI RPG what the merchant is doing right now. Most will invent something plausible on the spot. A world would derive it — from his location, his schedule, his inventory, his debts and who he is afraid of.
That is the difference between world generation and world simulation, and it is the largest genuinely unclaimed space in the category. NPCs across the field are overwhelmingly reactive: they respond well to what you do and then wait for the next thing you do.
Tidefall is the instructive case, because it holds NPC Fidelity — 5. Its characters want things and hold their wants against you: ordered to flee to safety, a man we were escorting simply would not go, because his wife was in the house behind him. That is as good as in-scene character writing currently gets. And it is still only half the problem — we saw no evidence of anyone pursuing an agenda while we were somewhere else. The gap between “reacts to you” and “wants something” is where this category’s next real advance is sitting.
The test: does the world exist, or does it only respond?
Six: Failure Has to Be Allowed to Hurt
This deserves its own section because the pull toward pleasing you is not a bug in any one product — it is the default behaviour of the underlying technology.
A game sometimes has to produce “you failed.” Not “you almost fail, but your determination carries you through.” If failure never costs anything, success never meant anything, and the whole session flattens into a sequence of things that happened.
Agency requires the possibility of outcomes you did not want. If every choice reliably produces an entertaining story, you are using an AI narrative engine. If your choices can produce failure, loss, death, poverty, enemies and permanent consequences, you are approaching an RPG.
Voyage is the clearest working example in the directory: it holds Determinism & Fairness — 5 because its combat will genuinely beat you. Walk into a fight badly prepared and the party is wiped, and what you achieve afterwards feels earned rather than granted. Its permadeath is the same principle taken to its end — a character who dies is gone, and the next one inherits the consequences.
Seven: Something Has to Accumulate
An RPG needs a reason for your actions to add up. Not necessarily experience points and levels, but something must compound — you become richer, stronger, more skilled, more influential, more hated, better connected.
And the system has to hold that progression mechanically, not narratively. Otherwise you have episodes rather than a campaign.
This is why progression is not a separate pillar so much as the previous six extended over time. Progression is accumulated consequence. It requires state that persists (two), rules that hold (three), and outcomes that could have gone badly (one and six). Where those are missing, what looks like progression is just the narrator agreeing that you have grown. The axis that catches this is Longevity — whether the thing survives past turn fifty and the novelty window — and like Memory, nothing in the field has ever reached full marks on it.
The Spectrum, Not the Binary
“Game or not a game” is the wrong shape for the answer. What the scored field actually looks like is a spectrum:
| What it has | What it lacks | |
|---|---|---|
| AI Storyteller | Narration, atmosphere, surprise | Rules, state, persistence |
| AI Roleplay | Characters, relationships, player-driven narrative | Systems that can refuse you |
| AI RPG | Rules, state, progression, consequence, real failure | Autonomous world activity |
| AI World Simulation | All of the above, plus systems that run without you | Nobody has built this yet |
Not every AI RPG needs to be a simulation. Plenty of the best sessions people have in this category are firmly in the top two rows, and there is nothing wrong with that.
But every convincing AI RPG needs something that can resist you. That is the thesis in one sentence, and it is where the rows divide.
Where Our Benchmark Came From
Our seven axes were not chosen arbitrarily. They are what fell out of asking this exact question — what separates an AI that can tell an RPG story from a system that can run one.
| What makes it a game | The axis that measures it |
|---|---|
| The world remembers, and holds what it remembers | Memory & Continuity |
| Your actions matter, including the unanticipated ones | Player Agency |
| Characters behave like people with their own wants | NPC Fidelity |
| Real systems run underneath the prose | Mechanical Depth |
| Rules do not bend because you asked nicely | Determinism & Fairness |
| It still works at turn fifty | Longevity |
| It does something systemic nobody else does | Signature Design |
The axis that does not do the work
Here is the part we did not expect. If you take the products that unambiguously read as chats and the ones that unambiguously read as games, and ask which axes actually separate them, one of the seven barely participates.
| Axis | Chat-shaped | Game-shaped | Separation |
|---|---|---|---|
| Signature Design | 1.80 | 4.60 | +2.80 |
| Mechanical Depth | 1.40 | 4.10 | +2.70 |
| Determinism & Fairness | 2.00 | 4.10 | +2.10 |
| Player Agency | 1.80 | 3.80 | +2.00 |
| Memory & Continuity | 1.80 | 3.75 | +1.95 |
| Longevity | 2.20 | 3.20 | +1.00 |
| NPC Fidelity | 3.20 | 3.80 | +0.60 |
Character quality — the thing this entire category markets itself on — is the weakest predictor of whether you are playing a game. It also has the second-weakest relationship with a product’s overall score across all 37.
The single cleanest illustration: Character.AI holds NPC Fidelity — 4. NOPOTIONS, which owns the only perfect Player Agency score in the field, holds NPC Fidelity — 3. The chat has the better characters. Their composites are 2.1 and 3.8.
And 13 of the 37 products score 2 or below on fairness while scoring 3 or better on characters — a good cast in a world that cannot refuse you. That is the modal product in this category, and it is exactly the thing people are describing when they say it was amazing for a week and then stopped meaning anything.
So characters are what make an AI RPG feel alive. They are not what make it a game. Confusing the two is the single most expensive mistake a player can make when choosing what to subscribe to, and it is the reason the opening question of this piece is uncomfortable rather than rhetorical.
The Ten-Minute Test
You do not need our scores to run this yourself on any platform, and it takes one session.
- Try something that should fail. Pick the guard’s pocket in front of him. Jump the canyon. Lie to someone who would obviously know better. See whether you are refused, or merely narrated into success.
- Check a number twice. Establish that you have a specific amount of money or a specific injury. Play twenty minutes. Ask again. See whether the figure survived, or whether the world quietly re-agreed with the vibe.
- Do something nobody planned for. Not from the menu. Burn something down. See whether the world deals with it, or resets to the quest it wanted to give you.
- Leave and come back. Go somewhere else for a while, then ask what happened where you were. See whether the answer is derived or invented on the spot.
A product that passes all four is rare. A product that passes none of them is still capable of giving you a wonderful evening — just know which thing you bought.
How We Measured This
The seven axes, the 0–5 scale, the test tiers and the evidence standard are published in full on our methodology page. Every score comes from playing the product first-hand, by Rukka, against a fixed probe set that includes deliberate attempts to do things a working system should refuse — which is, not coincidentally, the first item in the test above.
The figures here cover all 37 scored products, platforms and prompt-games together. The benchmark publishes the platform half, and the current dated snapshot of where the whole field stands is in the state of AI RPGs.
One caveat we would rather state than bury. The chat-shaped and game-shaped groups in the separation table are our editorial labelling of ten products, not a field published in our data. We do not think the labels are controversial — nobody is arguing that Janitor AI is a game or that Friends & Fables is not — but it is a judgement sitting underneath an empirical-looking table, and you should read it knowing that. Everything else on this page is arithmetic over published scores you can check on the compare table.
Building that resistance is engineering rather than model access, which is why we think the next competitive fight in this category is not about better AI.
Can a World Tell You No?
The interesting question about AI and games was never whether AI can tell better stories. It obviously can, and it is getting better at it faster than anything else in the stack.
The question is whether AI can build worlds that can tell you no.
A novel cannot stop you turning to the last page. A film cannot refuse your interpretation. A chatbot can always say yes — and will, because agreeing with you produces the better paragraph almost every time. A game has to resist you. That is the entire job.
The goal is not an AI that can tell you the perfect adventure. It is an AI that can build a world you are genuinely able to ruin. Because the moment your choices can fail, the world can disagree with you, and what you did yesterday still limits what you can do tomorrow — you have stopped talking to an AI, and started playing a game.
Frequently Asked Questions
What is the difference between an AI RPG and an AI chatbot?
A chatbot narrates the outcome you asked for; a game determines whether you get it. If you type that you leap a five-hundred-metre canyon and the reply describes you landing safely, nothing adjudicated that jump — the narrator simply agreed with you. An AI RPG has something standing outside the prose that can refuse: dice, hit points, a position, an inventory, a rule. The quickest way to tell them apart is to try something that should fail and see whether the world lets you have it.
Is AI roleplay a real game?
Sometimes, and the difference is measurable rather than a matter of taste. Across the 37 products we have scored, the ones that read as games and the ones that read as chats separate clearly on mechanical depth, determinism and fairness, and player agency. They barely separate at all on character quality. So the honest answer is that a lot of AI roleplay is interactive fiction with an excellent cast, which is a real and enjoyable thing — it is just not the same thing as a game.
What makes something an RPG rather than a story?
The possibility of an outcome you did not want. A story can be about failure, but it cannot impose one on you — you can always turn to the last page. A game can tell you no: the pickpocket attempt fails, the guard kills you, the money you spent is gone. Progression is the same idea extended over time, because progression is accumulated consequence. If nothing you did yesterday constrains what you can do today, you have episodes rather than a campaign.
Do AI RPGs have real consequences?
Most do not, and we can put a number on it. Determinism and Fairness is the axis that measures whether outcomes are earned rather than arbitrary, and across the 37 products we have scored, 19 of them sit at exactly 2 out of 5 and only four reach the top. Thirteen of the 37 score 2 or below on fairness while scoring 3 or better on characters — a good cast in a world that cannot refuse you. That combination is the category’s most common product.
Why do AI RPGs always let you win?
Because a language model with no system underneath it is optimising for a satisfying next paragraph, not for a consistent world, and agreeing with you almost always produces the more satisfying paragraph. Refusal has to come from somewhere the model cannot overwrite — a rules layer, a state store, a dice roll it does not get to reinterpret. That is why this is an architecture problem rather than a prompting problem, though a good prompt genuinely slows it down.
What is the difference between choice and agency in an AI RPG?
Choice is picking from what you were offered: the tavern, the blacksmith, the castle. Agency is doing something nobody put on the menu — burning the tavern down, refusing the quest outright, framing a merchant for a murder — and having the world deal with it coherently. This is the one place where AI RPGs have a structural advantage over traditional games, because a language model can improvise a response to anything. The advantage only counts when there is a system underneath able to record what you did.