Your AI Dungeon Master Is Too Generous (and Why That Ruins It)

If your AI dungeon master has never really let you fail, then nothing you have done in it counts — and some part of you already knows that, which is usually why the campaign quietly stopped being interesting around hour three.

This is the most common complaint we hear about AI RPGs that are otherwise going well. The writing is good. The characters are good. And yet the whole thing feels weightless, because every plan works, every guard is convinced, and every wound turns out to be survivable. We scored this. Across 28 products measured on seven axes from first-hand play, Determinism & Fairness — whether outcomes are earned and consistent rather than arbitrary and luck-washed — averages 2.39 out of 5.

The shape of that distribution is stranger than the average:

Determinism & FairnessProducts
51
40
310
215
12

Nothing scores 4. One product clears 3, and then there is a cliff. Twenty-five of 28 sit at either 2 or 3 — which is to say the overwhelming majority of AI RPGs are somewhere between “occasionally arbitrary” and “will agree with almost anything you say confidently.”

Why Models Default to Yes

The generosity isn’t a bug any single developer introduced. It comes from what the underlying model is for.

A language model produces the most plausible continuation of the text so far. When that text is “I convince the guard that I’m the Duke’s cousin,” the most plausible next passage — by an enormous margin — is the one where the guard believes you. Refusal is rarer in the training data, harder to write well, and reads as unhelpful. Everything about the model’s disposition points toward yes.

Then there’s the second problem, which is the one that actually decides it: there’s usually nothing underneath to say no from. A refusal needs a reason — a difficulty number, a guard who has been told you look nothing like the Duke, a record of your reputation in this district. Most AI RPGs have none of that. We measured this too: Mechanical Depth is the field’s weakest axis, with 8 of 28 products having effectively no systems at all. A game with no state to check against can only decide outcomes on vibes, and the vibe is nearly always accommodating.

So the failure compounds. The model wants to say yes, and there is nothing in the architecture positioned to overrule it.

The Four Tells

You can diagnose this in about ten minutes.

Confidence works. Phrase a request with certainty and it succeeds; phrase the same request hesitantly and it also succeeds, but with hedging prose. If how you asked changed the outcome, nothing was being adjudicated.

Failure is always narrative, never mechanical. You do sometimes fail — but only when failing makes for a better scene. You never fail because a number said so. The difference matters: one is a story beat the game chose, the other is a result you could have anticipated.

Resources never bite. You have gold, supplies, ammunition, spell slots. Nothing ever runs out at an inconvenient moment. Scarcity is described but never enforced.

No plan is ever simply wrong. The most reliable test. Propose something genuinely bad — a plan with an obvious flaw a competent GM would let you walk into. If the world reshapes to make it work anyway, you are not playing a game. You are co-writing a story that flatters you, which is a fine thing to want and a very different product from the one most of these are sold as.

The Excuse That Doesn’t Survive the Data

There is a standard defence of restrictive AI RPGs: of course we limit what the player can attempt — that’s how we keep outcomes consistent. Freedom and fairness are assumed to trade against each other. Let the player do anything and you can’t possibly adjudicate it properly.

Our numbers do not support that.

Across all 28 scored products, Player Agency and Determinism & Fairness correlate positively — the products that give you more freedom tend to be the fairer ones, not the looser ones. Products scoring 3 or better on fairness average 2.64 on Player Agency. Products scoring 2 or below average 2.00.

And the sharpest version of the point is at the bottom of the agency axis:

ProductPlayer AgencyDeterminism & Fairness
Dunia01
Hidden Door12
WyrdTale12

These are the three lowest agency scores in the directory, and none of them is fair. Dunia scores the only Player Agency — 0 we have awarded, for refusing to let your character take certain actions at all.

Note that they get there by opposite routes, which is the part that matters. Dunia and Hidden Door are genuinely restrictive — they narrow what you may attempt. WyrdTale restricts almost nothing; it simply keeps speaking and acting as your character, which costs you authorship just as completely without ever telling you no. Three products, two different mechanisms, and the fairness scores land in the same place regardless.

Restricting the player does not buy consistent outcomes. It buys a smaller game that is still arbitrary inside its smaller boundaries. If constraint alone produced fairness, the restrictive pair here would be the fairest products on the board, and they are close to the opposite — while the permissive one sitting beside them is no fairer.

A caveat, stated plainly because the number deserves it: 28 products scored 0–5 by one tester is a directional finding, not a statistical proof. The correlation is moderate (r ≈ 0.49). What it is strong enough to do is kill the excuse — whatever explains fairness, it clearly isn’t restriction.

What the One Exception Actually Does

Friends & Fables is the single product above 3, at Determinism & Fairness — 5, and what earns it is precisely the thing the trade-off theory says is impossible: it holds both behaviours at once.

You cannot talk your way into an unearned reward. Phrasing a request confidently doesn’t produce a success. But refusing to be prompt-hacked has not been implemented as refusing everything — creativity and genuine problem-solving get rewarded rather than blocked, and an unorthodox approach is allowed to be tried on its merits.

Those two normally trade off. Most products that can’t be talked into things achieve it by being unable to be talked into anything, which is a different failure wearing a stricter coat. Holding both is rare enough that we called it the platform’s most underrated quality in the full review, and it is not a coincidence that the same product scores highest on Mechanical Depth. Fairness is downstream of having something to be fair with.

The Worked Failure: Generosity as a Design Choice

The clearest example of the opposite runs in the other direction — not arbitrary, just relentlessly kind.

Valkyrie’s Biggest Gig scores Determinism & Fairness — 1. The game itself is competent and its central idea is genuinely good: it’s a cyberpunk RPG where limb loss and cybernetic augmentation are the point, so your body becomes a resource you spend. That premise only works if spending hurts. Instead the reward economy is generous enough that almost nothing has to be earned, which quietly removes the tension the whole design was built to create.

It’s the most instructive kind of failure, because nothing is broken. Every part works. The game simply never charges you for anything, and a cyberpunk story where nothing costs you is a tourism brochure.

How to Put the Stakes Back

You can’t add a rules engine to a product that doesn’t have one. You can stop the model resolving everything in your favour.

  • Define failure before you play. In your opening message, state what losing looks like: what can kill you, what can be permanently lost, what the world does if you fail. Models are far better at honouring a failure condition they were given than at inventing one mid-scene.
  • Ask for the result before the prose. Require the narrator to state the outcome first, then describe it. This is a small change with a large effect, because it stops the description from deciding the result retroactively.
  • Ask what it would take to fail. Before a risky action, ask the game what would make this go wrong. Then hold it to its own answer. Models are much more willing to enforce a standard they just articulated.
  • Keep resources where you can see them. A number the game must reconcile against is harder to quietly inflate. Our free campaign memory tool is built for keeping that record outside the chat.
  • Don’t confuse “harder” with “grimmer.” Asking a model to be more difficult usually produces darker descriptions of identical outcomes. Ask for consequences, not atmosphere.

If fairness is the thing you care about most, rank the directory by that axis directly rather than by overall score — and note how little the top of that list resembles the top of the composite one.

Fairness is one of seven problems an AI RPG has to solve, and the one with the strangest distribution — the seven problems, ranked by how badly the field handles them puts it in context.

How We Measured This

Determinism & Fairness is one of seven axes in methodology v1.0, scored 0–5 from first-hand play by Rukka, with a fixed probe set run against every entry. The fairness probes deliberately include attempts that should fail: unearned persuasion, plans with obvious flaws, and requests phrased with unjustified confidence.

The correlation reported above compares Player Agency and Determinism & Fairness across all 28 scored products. It is a moderate positive relationship (r ≈ 0.49) on a small, ordinal, single-tester dataset — enough to contradict the claim that restriction produces fairness, not enough to prove what does.

Two limits worth naming. A score of 1 here means outcomes are arbitrary or unearned, not that the product is unpleasant to play — Valkyrie’s Biggest Gig is a 1 on this axis and still worth a session. And we don’t score pre-release software, so this reflects what shipped.

The figures here cover all 28 scored products, platforms and prompt-games together. The Arcanum AI RPG Benchmark publishes the platform half — the 16 scored platforms, every axis — where the per-axis averages run slightly higher. If a number here doesn’t match your experience of the same product, tell us: [email protected].

Frequently Asked Questions

Why is my AI dungeon master so easy? Because a language model is trained to produce a helpful, satisfying continuation, and refusing you rarely looks like either. When you describe a clever plan, the most plausible next passage is the one where the plan works — so it works. There is usually no number underneath deciding otherwise, and no record of how difficult the world was supposed to be. The result is a game that agrees with you, which feels like success for about an hour and then stops meaning anything.

How do I make an AI RPG harder? Ask for consequences rather than difficulty. Telling a model to be harder usually produces grimmer descriptions of the same outcomes. Instead define what failure looks like before you play, keep a visible record of resources that can actually run out, and require the narrator to state a result before describing it. The most effective single habit is asking what it would take to fail at something, then holding the game to its own answer.

Which AI RPG actually has real stakes? Friends & Fables is the only product in our directory scoring above 3 on Determinism & Fairness, at 5. What earns it is holding two behaviours that normally trade off: you cannot talk your way into an unearned reward, and unorthodox approaches are still allowed to be tried on their merits. Most products that refuse to be talked into things achieve it by refusing almost everything, which is a different failure rather than a fix.

Does a stricter AI RPG mean a fairer one? No, and our data points the other way. Across 28 scored products, Player Agency and Determinism & Fairness correlate positively rather than negatively — the products that give you the most freedom tend to be the fairer ones. The three lowest-agency products in our directory score 1, 2 and 2 on fairness. Restricting what a player may attempt does not, on this evidence, buy consistent outcomes.

Is an AI RPG that says no just broken? It depends on what it is refusing. Refusing an unearned success is the system working. Refusing an action your character should simply be able to attempt is a different thing entirely, and we score it on Player Agency rather than fairness — one product in our directory scores 0 there for exactly this. A good AI RPG lets you try almost anything and does not guarantee it works.