Why AI RPGs Can't Do Mechanics (and the Two That Can)
Mechanical depth is the difference between a game that runs systems and a game that describes them — and it is the single thing AI RPGs are worst at.
We now score every product in our directory on seven axes, each from first-hand play. Across 28 scored products — 16 platforms and 12 prompt-games — Mechanical Depth averages 2.00 out of 5 — the lowest of the seven by a clear margin. Eight of the 28 score 1 or below, meaning there is essentially no system underneath the story at all. Only two score above 3.
That is not a complaint about polish. It is a structural fact about what language models are, and it explains a failure most players have felt without being able to name: the moment when your inventory quietly stops mattering, your injuries heal because the scene changed tone, and the difficulty of a fight turns out to depend on how dramatic the writing felt rather than on any number you could check.
What Mechanical Depth Actually Measures
The axis asks one question: are there real systems underneath, or is this narration wearing dice?
A mechanically deep game resolves a fight using values that exist before the sentence describing it is written. Your armour class is a number. Cover is a modifier. The orc’s resistance to slashing damage is a property it had when it entered the scene. The prose is generated from the outcome.
A shallow game does the reverse. It writes a convincing account of a fight, and any numbers appear inside that account as decoration. The sentence “you roll a 17 and cleave through its guard” is not evidence that anything was rolled. It is evidence that the model knows what such a sentence looks like.
Both can produce excellent reading. Only one produces outcomes you could have predicted beforehand or disputed afterwards, and that distinction is the whole axis. It is also why depth is not the same as complexity: a game can present character sheets, spell slots and encumbrance tables and still score low, because presenting a system and enforcing one are different things.
The Numbers: The Field’s Worst Axis
Here is the whole field, scored on the same rubric:
| Mechanical Depth | Products | Share |
|---|---|---|
| 5 | 1 | 4% |
| 4 | 1 | 4% |
| 3 | 4 | 14% |
| 2 | 14 | 50% |
| 1 | 7 | 25% |
| 0 | 1 | 4% |
Three things stand out.
The middle is enormous. Fourteen of 28 products — half the field — score exactly 2. That is the score for a game that gestures at systems: it has an inventory you can ask about, it will mention hit points, and none of it holds up under pressure.
The bottom is crowded. Eight products score 1 or below. For comparison, NPC Fidelity — the field’s strongest axis, averaging 2.86 — has exactly one product that low. The category learned to write characters long before it learned to run games.
The top is nearly empty. Two products above 3, out of 28. Whatever is hard here, almost nobody has solved it.
Why a Language Model Cannot Hold a System
The reason is not laziness on the part of developers, and it is not a prompt that could have been written better. It is architectural, and it has three parts.
A model has nowhere to put a number. Between one turn and the next, the only thing carrying your gold total forward is the conversation itself. There is no variable. There is no row in a table. The number persists exactly as long as the text mentioning it stays within reach of the model’s attention, and no longer.
Plausibility beats arithmetic. A language model generates what fits. If the scene has become desperate, a desperate amount of remaining health fits, whether or not it matches the number from twenty turns ago. The model is not calculating incorrectly — it is not calculating. It is producing the most likely continuation, and “most likely” is a literary judgement rather than a mathematical one.
Recency wins. As a session grows, early facts sit further from the model’s attention than recent ones. The system you established in your opening message is competing with four hours of subsequent narration, and it loses gradually. This is the same mechanism behind campaigns falling apart in general, which we covered in why your AI campaign falls apart at turn 50.
Put together: a prompt is not a database. Everything a prompt establishes lives in the same finite, shifting window as everything else, so a rules layer built purely in the prompt degrades on exactly the same curve as the plot.
The Four Ways It Fails in Play
Recognising the failure is more useful than knowing the theory, and it always takes one of four shapes.
Silent stat drift. Your health, gold or ammunition changes without an event causing it. The tell is that it never drifts against you in a way that creates a problem — drift resolves toward whatever the current scene needs.
Inventory amnesia. You are carrying something the story has stopped modelling. Try using an item you acquired long ago and rarely mention: if it works exactly as well as it would have when new, and nobody accounts for how many you have left, there is no inventory — there is a memory of having mentioned one.
Mood-based difficulty. The hardness of a fight tracks the emotional temperature of the writing rather than any statistic. Bosses become dangerous when the prose turns ominous and become manageable when the scene wants to resolve. This is where a missing mechanical layer shows up as a missing sense of stakes — covered separately in why your AI dungeon master is too generous.
Rules that apply when remembered. The clearest sign. A rule enforced strictly in session one is quietly ignored in session four — not overruled, not argued, just absent. This is the one that separates a real system from a described one, because a real system cannot forget to run.
The First Exception: Friends & Fables
Friends & Fables is the only product in the directory scoring Mechanical Depth — 5, and it earns it in the place the axis is hardest: combat.
Its V3 combat system runs on hex battlemaps with line-of-sight calculations, cover bonuses, resistances and vulnerabilities. The important part is not the feature list — plenty of products list features. It is that these resolve outside the narration. Position is a coordinate. Cover either applies or it doesn’t. The description of the swing is written after the question of whether it lands has already been settled somewhere the model does not get a vote.
That is the entire trick, and it is an engineering decision rather than a writing one. The full review covers where the same complexity costs it elsewhere — a system this heavy can force encounters that a lighter game would have let you talk your way out of.
The Second Exception, and the Trap Inside It
AI Realm scores Mechanical Depth — 4 on a genuinely built-out D&D 5e layer: 12 official classes, 9 races, a 27-point-buy ability system, rules applied rather than described. On this axis alone it is the second-best thing in the category.
It also scores Memory & Continuity — 1, the joint-lowest in the directory.
That combination is the most instructive result in the entire dataset, because it isolates the variable. AI Realm built the rules and did not build the place to keep them. It uses a standard context window rather than an external state database, so the mechanical reliability that is genuinely impressive in session one degrades by session four: spell rules enforced correctly early get quietly mishandled later, and plot threads reset. Our review describes the late game as a hard ceiling that no interface feature resolves.
So AI Realm is not a counter-example to the thesis. It is the proof of it. Rules are not the hard part. Somewhere to keep them is.
The Design Lesson: Depth Without State Decays
Line the two exceptions up and the principle is unambiguous.
| Mechanical Depth | Memory & Continuity | What it means | |
|---|---|---|---|
| Friends & Fables | 5 | 3.5 | Systems resolve outside the model, and survive |
| AI Realm | 4 | 1 | Systems resolve outside the model, and are forgotten |
Friends & Fables pairs the directory’s highest Mechanical Depth with a Memory score behind only Hidden Door and WyrdTale — the first scores Mechanical Depth — 1 and is not trying to run systems at all, and the second scores 3. Among products that run deep systems, nothing remembers them better. That is not a coincidence and it is not two separate achievements. A mechanical layer is only as durable as the state layer beneath it, because a rule the game cannot remember applying is indistinguishable, from the player’s side, from a rule that was never there.
The corollary is the useful part for anyone building in this space: adding rules to a system with no persistent state does not raise mechanical depth for long. It raises it for a session. The measurable ceiling is set by where the numbers live, not by how many of them there are.
What This Means If You’re Playing
You cannot install a state layer into a product that lacks one, but you can stop relying on the one it pretends to have.
- Test before you invest. In your first twenty minutes, attempt something you should plainly fail. If the world bends, there is no system underneath and no amount of later effort will create one.
- Keep the authoritative copy outside the chat. The single highest-value habit in AI roleplay. Maintain your own record of inventory, injuries and standing, and paste it back in when things wobble — our free campaign memory tool exists for exactly this.
- Prefer tags to numbers. “Wounded left arm, cannot hold a shield” survives summarisation far better than “HP 14/30”, because it carries its own consequence and reads as narrative rather than bookkeeping.
- Restate the system periodically. Not because the model forgot in the human sense, but because restating moves the rules back into recent context where the model actually looks.
The full technique set is in our guide to making stats and system windows actually stick. If mechanics matter more to you than anything else, rank the directory by Mechanical Depth directly rather than by overall score — the top of that list is a very different set of products from the top of the composite one, which is the point of scoring axes separately at all.
Mechanics are one of seven problems an AI RPG has to solve, and by our measurements the hardest — the seven problems, ranked by how badly the field handles them covers the rest and what they share.
How We Measured This
Every number here comes from playing the product, not from reading its marketing. Mechanical Depth is one of seven axes in methodology v1.0, scored 0–5, with a fixed probe set run against each entry — including deliberately attempting actions that a real system should refuse.
The figures on this page cover all 28 scored products — 16 platforms and 12 prompt-games — because the question here is what a prompt can and can’t hold, and the prompt-games are the cleanest evidence of that. The Arcanum AI RPG Benchmark publishes the platform half, all 16, every axis; its per-axis averages run slightly higher for that reason. Scores are from first-hand play by Rukka, and every rated review states which test tier it received and what it previously scored if the number has changed.
Two limits worth stating. A 0 on this axis means the capability is absent, not that the product is bad — a one-line prompt scoring 0 is being measured against a job it never applied for. And pre-release software is not scored at all, so the board reflects what shipped, not what is coming.
If a number here doesn’t survive contact with your experience of the same product, we want to hear it: [email protected]. Disagreeing with a verdict isn’t a correction; a score whose reasoning doesn’t hold up is.
Frequently Asked Questions
What is mechanical depth in an AI RPG? Mechanical depth is whether real systems run underneath the story, or whether the game is narration wearing dice. A mechanically deep game resolves a fight using numbers that exist before the sentence is written — hit points, cover, resistances, an inventory that is a list rather than a memory. A shallow one writes a convincing description of a fight and derives the numbers from the description afterwards. Both can read well. Only one produces outcomes you could have predicted or argued with.
Why can’t AI RPGs handle combat and stats? Because a language model predicts plausible text, and it has no place to keep a number between turns except the conversation itself. As a session grows, older facts sit further from the model’s attention, so your gold total, your inventory and your injuries drift toward whatever sounds reasonable now rather than what was true forty turns ago. This is the same underlying limit that causes campaigns to lose continuity, and no amount of prompt discipline removes it — a prompt is not a database.
Which AI RPG has the best mechanics? Friends & Fables, which is the only product in our directory scoring 5 on Mechanical Depth. It runs V3 hex-battlemap combat with line-of-sight calculations, cover bonuses, resistances and vulnerabilities — systems that resolve outside the narration rather than inside it. AI Realm is second at 4, with a genuinely built-out D&D 5e layer, but it pairs that with the joint-lowest Memory score in the directory, so the rules run a campaign the platform cannot remember.
Do AI RPGs actually roll dice? Some do and most don’t. A platform with a real mechanical layer generates a number, applies modifiers and then narrates the result of that number. A platform without one writes the narration first and mentions a roll inside it, which means the roll never constrained anything. You can usually tell within a few turns by attempting something you should clearly fail: if the story bends to accommodate you, there was no roll underneath it.
Can a better prompt fix stat tracking? It helps and it does not solve it. Structured status blocks, tag-based state and restating the system periodically all slow the drift measurably. What they cannot do is create an authoritative record, because everything the prompt establishes still lives in the same context window that is filling up. The reliable fix is to keep the authoritative copy outside the chat and paste it back in, which is a workaround for a missing state layer rather than a replacement for one.