Top

Why AI RPGs Can't Do Mechanics (and the Two That Can)

Mechanical depth is the difference between a game that runs systems and a game that describes them — and it is the single thing AI RPGs are worst at.

We now score every product in our directory on seven axes, each from first-hand play. Across 37 scored products — 25 platforms and 12 prompt-games — Mechanical Depth averages 2.27 out of 5 — the lowest of the seven by a clear margin. Nine of the 37 score 1 or below, meaning there is essentially no system underneath the story at all. Only six score above 3.

That is not a complaint about polish. It is a structural fact about what language models are, and it explains a failure most players have felt without being able to name: the moment when your inventory quietly stops mattering, your injuries heal because the scene changed tone, and the difficulty of a fight turns out to depend on how dramatic the writing felt rather than on any number you could check.

What Mechanical Depth Actually Measures

If the prior question — whether the thing in front of you is a game at all — is the one you actually have, what makes an AI RPG feel like a game draws that line first and comes back here for the mechanism.

The axis asks one question: are there real systems underneath, or is this narration wearing dice?

A mechanically deep game resolves a fight using values that exist before the sentence describing it is written. Your armour class is a number. Cover is a modifier. The orc’s resistance to slashing damage is a property it had when it entered the scene. The prose is generated from the outcome.

A shallow game does the reverse. It writes a convincing account of a fight, and any numbers appear inside that account as decoration. The sentence “you roll a 17 and cleave through its guard” is not evidence that anything was rolled. It is evidence that the model knows what such a sentence looks like.

Both can produce excellent reading. Only one produces outcomes you could have predicted beforehand or disputed afterwards, and that distinction is the whole axis. It is also why depth is not the same as complexity: a game can present character sheets, spell slots and encumbrance tables and still score low, because presenting a system and enforcing one are different things.

The Numbers: The Field’s Worst Axis

Here is the whole field, scored on the same rubric:

Mechanical DepthProductsShare
525%
425%
3.525%
3719%
21541%
1822%
013%

Three things stand out.

The middle is enormous. Fifteen of 37 products — more than four in ten — score exactly 2. That is the score for a game that gestures at systems: it has an inventory you can ask about, it will mention hit points, and none of it holds up under pressure.

The bottom is crowded. Nine products score 1 or below. For comparison, NPC Fidelity — the field’s strongest axis, averaging 2.97 — has exactly one product that low. The category learned to write characters long before it learned to run games.

One of those nine is worth separating from the rest. Tidefall is marketed as built on the D&D SRD and scores Mechanical Depth — 1 anyway: no inventory screen, no currency display, no quest journal, and combat that reports the opposition as “a swarm” rather than as a number of creatures. Its world is dense and hand-made and its encounters are designed by people — what is missing is any layer letting you operate on them. The other eight are thin because nobody built the systems. This one is thin in front of a world that would reward them, which is a different and more fixable problem.

The top is nearly empty. Six products above 3, out of 37. It is the field’s worst axis by average and, oddly, one of the five where somebody has still reached full marks — a distinction our state-of-the-field edition unpacks. Whatever is hard here, almost nobody has solved it.

Why a Language Model Cannot Hold a System

The reason is not laziness on the part of developers, and it is not a prompt that could have been written better. It is architectural, and it has three parts.

A model has nowhere to put a number. Between one turn and the next, the only thing carrying your gold total forward is the conversation itself. There is no variable. There is no row in a table. The number persists exactly as long as the text mentioning it stays within reach of the model’s attention, and no longer.

Plausibility beats arithmetic. A language model generates what fits. If the scene has become desperate, a desperate amount of remaining health fits, whether or not it matches the number from twenty turns ago. The model is not calculating incorrectly — it is not calculating. It is producing the most likely continuation, and “most likely” is a literary judgement rather than a mathematical one.

Recency wins. As a session grows, early facts sit further from the model’s attention than recent ones. The system you established in your opening message is competing with four hours of subsequent narration, and it loses gradually. This is the same mechanism behind campaigns falling apart in general, which we covered in why your AI campaign falls apart at turn 50.

Put together: a prompt is not a database. Everything a prompt establishes lives in the same finite, shifting window as everything else, so a rules layer built purely in the prompt degrades on exactly the same curve as the plot.

The Four Ways It Fails in Play

Recognising the failure is more useful than knowing the theory, and it always takes one of four shapes.

Silent stat drift. Your health, gold or ammunition changes without an event causing it. The tell is that it never drifts against you in a way that creates a problem — drift resolves toward whatever the current scene needs.

Inventory amnesia. You are carrying something the story has stopped modelling. Try using an item you acquired long ago and rarely mention: if it works exactly as well as it would have when new, and nobody accounts for how many you have left, there is no inventory — there is a memory of having mentioned one.

Mood-based difficulty. The hardness of a fight tracks the emotional temperature of the writing rather than any statistic. Bosses become dangerous when the prose turns ominous and become manageable when the scene wants to resolve. This is where a missing mechanical layer shows up as a missing sense of stakes — covered separately in why your AI dungeon master is too generous.

Rules that apply when remembered. The clearest sign. A rule enforced strictly in session one is quietly ignored in session four — not overruled, not argued, just absent. This is the one that separates a real system from a described one, because a real system cannot forget to run.

The First Exception: Friends & Fables

Friends & Fables is one of two products in the directory scoring Mechanical Depth — 5, and it earns it in the place the axis is hardest: combat.

Its V3 combat system runs on hex battlemaps with line-of-sight calculations, cover bonuses, resistances and vulnerabilities. The important part is not the feature list — plenty of products list features. It is that these resolve outside the narration. Position is a coordinate. Cover either applies or it doesn’t. The description of the swing is written after the question of whether it lands has already been settled somewhere the model does not get a vote.

That is the entire trick, and it is an engineering decision rather than a writing one. The full review covers where the same complexity costs it elsewhere — a system this heavy can force encounters that a lighter game would have let you talk your way out of.

The Second Exception, and the Trap Inside It

AI Realm scores Mechanical Depth — 4 on a genuinely built-out D&D 5e layer: 12 official classes, 9 races, a 27-point-buy ability system, rules applied rather than described. On this axis alone only the two products at 5 beat it, and only Voyage matches it.

It also scores Memory & Continuity — 1, the joint-lowest in the directory.

That combination is the most instructive result in the entire dataset, because it isolates the variable. AI Realm built the rules and did not build the place to keep them. It uses a standard context window rather than an external state database, so the mechanical reliability that is genuinely impressive in session one degrades by session four: spell rules enforced correctly early get quietly mishandled later, and plot threads reset. Our review describes the late game as a hard ceiling that no interface feature resolves.

So AI Realm is not a counter-example to the thesis. It is the proof of it. Rules are not the hard part. Somewhere to keep them is.

The Third Exception: Rules Somebody Else Wrote

Craft scores Mechanical Depth — 5 without shipping a ruleset of its own at all.

The Friends & Fables studio’s successor platform is system-agnostic by design: it hands the ruleset to the worldbuilder, who can write new gameplay systems as code rather than assembling them from a menu of platform features. What the platform guarantees is execution. In our full review the mechanics resolved exactly as the world’s author wrote them, every time — which is also why it takes Determinism & Fairness — 5, one of only four products to reach that number.

Read against the thesis, Craft is the same finding one layer up. The rules are not in the prompt; they are in an authored layer the model consults and cannot overrule. It does not matter that a creator wrote them rather than the studio. What matters is that they live somewhere the context window cannot erode.

The catch is the one no other exception has: Craft’s mechanical ceiling is a property of the world you load, not of the platform. A carefully built world plays deeper than anything else in this directory. A careless one plays like nothing at all, and the platform will run that faithfully too.

The Fourth Exception: The One That Says It Out Loud

NOPOTIONS scores Mechanical Depth — 3.5, and it is the only product in this directory whose developers describe the thesis of this page as their own architecture, in their own marketing: “the Narrator handles description and NPC behaviour, and the game engine handles the d20, the difficulty class, and the state change.”

That is the split, stated plainly. The engine rolls every check in the open — the natural die, the difficulty class and your modifiers shown separately — and the result stands whatever the prose would have preferred. Underneath it there is loot with rarity tiers and affixes that raise real values while worn, materials you harvest and craft, and 21 skills that level from what you actually did. None of it is narrated into existence.

It scores 3.5 rather than 5 for reasons that have nothing to do with the architecture: parts of the ruleset are not built yet. Combat has no flanking and no disadvantage, magic and most classes are visible placeholders, and there is no way to change the rules the way Craft lets you. It is the right design with a third of the content still to come, which is the most encouraging failure mode on this axis. The full review has the axis-by-axis detail.

It is also a useful correction to a lazy reading of this page. NOPOTIONS scores Memory & Continuity — 3 — mid-field — while holding depth above the field. Mechanical state and narrative state are not the same store, and this product keeps the first one better than the second: our inventory and skills never drifted, and the story did, losing track of what we were supposed to be doing.

The Design Lesson: Depth Without State Decays

This is also why we expect the next few years of this category to be won on engineering rather than on model access — the argument, with the scores behind it.

Line the four exceptions up and the principle is unambiguous.

Mechanical DepthMemory & ContinuityWhat it means
Craft54Systems are authored outside the model, by the worldbuilder
Friends & Fables53.5Systems resolve outside the model, and survive
NOPOTIONS3.53Systems resolve outside the model; the story layer drifts anyway
AI Realm41Systems resolve outside the model, and are forgotten

The two products at the top of this axis are also two of the best in the directory at holding on to what they run. Craft scores Memory & Continuity — 4, beaten only by Voyage’s 4.25. Friends & Fables scores 3.5, behind only Voyage, Craft, Hidden Door, WyrdTale and Tabled — and the other three run far shallower systems than it does. Among products that run deep systems, these two remember them best. That is not a coincidence and it is not two separate achievements. A mechanical layer is only as durable as the state layer beneath it, because a rule the game cannot remember applying is indistinguishable, from the player’s side, from a rule that was never there.

The corollary is the useful part for anyone building in this space: adding rules to a system with no persistent state does not raise mechanical depth for long. It raises it for a session. The measurable ceiling is set by where the numbers live, not by how many of them there are.

One limit on the principle is worth stating. Moving the dice out of the model settles whether a roll is honest; it does not settle when to call for one. In our first look at WorldAI, a free D&D 5e game still in development, the dice are rolled server-side and can be audited — and the game still asked us to roll just to summon our own lieutenants for a report. Deciding that an action is uncertain enough to test is a game master’s judgement, and no dice server makes it for you.

What This Means If You’re Playing

You cannot install a state layer into a product that lacks one, but you can stop relying on the one it pretends to have. For the genre where that reliance breaks fastest, and a session setup that works around it, see AI survival RPGs. (There is one architecture that lets you bring your own instead of settling for what a platform provides — see running an AI RPG on an MCP engine.)

  • Test before you invest. In your first twenty minutes, attempt something you should plainly fail. If the world bends, there is no system underneath and no amount of later effort will create one. The full sequence — five checks, one per failure this medium is prone to — is in how to test an AI RPG in your first twenty minutes.
  • Keep the authoritative copy outside the chat. The single highest-value habit in AI roleplay. Maintain your own record of inventory, injuries and standing, and paste it back in when things wobble — our free campaign memory tool exists for exactly this.
  • Prefer tags to numbers. “Wounded left arm, cannot hold a shield” survives summarisation far better than “HP 14/30”, because it carries its own consequence and reads as narrative rather than bookkeeping.
  • Restate the system periodically. Not because the model forgot in the human sense, but because restating moves the rules back into recent context where the model actually looks.

The full technique set is in our guide to making stats and system windows actually stick. If mechanics matter more to you than anything else, rank the directory by Mechanical Depth directly rather than by overall score — the top of that list is a very different set of products from the top of the composite one, which is the point of scoring axes separately at all.

Mechanics are one of seven problems an AI RPG has to solve, and by our measurements the hardest — the seven problems, ranked by how badly the field handles them covers the rest and what they share.

How We Measured This

Every number here comes from playing the product, not from reading its marketing. Mechanical Depth is one of seven axes in methodology v1.0, scored 0–5, with a fixed probe set run against each entry — including deliberately attempting actions that a real system should refuse.

The figures on this page cover all 37 scored products — 25 platforms and 12 prompt-games — because the question here is what a prompt can and can’t hold, and the prompt-games are the cleanest evidence of that. The Arcanum AI RPG Benchmark publishes the platform half, all 25, every axis; its per-axis averages run slightly higher for that reason. Scores are from first-hand play by Rukka, and every rated review states which test tier it received and what it previously scored if the number has changed.

Two limits worth stating. A 0 on this axis means the capability is absent, not that the product is bad — a one-line prompt scoring 0 is being measured against a job it never applied for. And pre-release software is not scored at all, so the board reflects what shipped, not what is coming.

If a number here doesn’t survive contact with your experience of the same product, we want to hear it: [email protected]. Disagreeing with a verdict isn’t a correction; a score whose reasoning doesn’t hold up is.

Frequently Asked Questions

What is mechanical depth in an AI RPG? Mechanical depth is whether real systems run underneath the story, or whether the game is narration wearing dice. A mechanically deep game resolves a fight using numbers that exist before the sentence is written — hit points, cover, resistances, an inventory that is a list rather than a memory. A shallow one writes a convincing description of a fight and derives the numbers from the description afterwards. Both can read well. Only one produces outcomes you could have predicted or argued with.

Why can’t AI RPGs handle combat and stats? Because a language model predicts plausible text, and it has no place to keep a number between turns except the conversation itself. As a session grows, older facts sit further from the model’s attention, so your gold total, your inventory and your injuries drift toward whatever sounds reasonable now rather than what was true forty turns ago. This is the same underlying limit that causes campaigns to lose continuity, and no amount of prompt discipline removes it — a prompt is not a database.

Which AI RPG has the best mechanics? Two products score 5 on Mechanical Depth, and they get there by opposite routes. Friends & Fables runs V3 hex-battlemap combat with line-of-sight calculations, cover bonuses, resistances and vulnerabilities — systems that resolve outside the narration rather than inside it. Craft ships no ruleset of its own and instead lets worldbuilders write new gameplay systems as code, so its ceiling is set by whoever built the world you loaded. AI Realm is third at 4, with a genuinely built-out D&D 5e layer, but it pairs that with the joint-lowest Memory score in the directory, so the rules run a campaign the platform cannot remember.

Do AI RPGs actually roll dice? Some do and most don’t. A platform with a real mechanical layer generates a number, applies modifiers and then narrates the result of that number. A platform without one writes the narration first and mentions a roll inside it, which means the roll never constrained anything. You can usually tell within a few turns by attempting something you should clearly fail: if the story bends to accommodate you, there was no roll underneath it.

Can a better prompt fix stat tracking? It helps and it does not solve it. Structured status blocks, tag-based state and restating the system periodically all slow the drift measurably. What they cannot do is create an authoritative record, because everything the prompt establishes still lives in the same context window that is filling up. The reliable fix is to keep the authoritative copy outside the chat and paste it back in, which is a workaround for a missing state layer rather than a replacement for one.