Top

Your AI Dungeon Master Is Too Generous (and Why That Ruins It)

If your AI dungeon master has never really let you fail, then nothing you have done in it counts — and some part of you already knows that, which is usually why the campaign quietly stopped being interesting around hour three.

This is the most common complaint we hear about AI RPGs that are otherwise going well. The writing is good. The characters are good. And yet the whole thing feels weightless, because every plan works, every guard is convinced, and every wound turns out to be survivable. We scored this. Across 37 products measured on seven axes from first-hand play, Determinism & Fairness — whether outcomes are earned and consistent rather than arbitrary and luck-washed — averages 2.61 out of 5.

The shape of that distribution is stranger than the average:

Determinism & FairnessProducts
54
40
3.51
311
219
12

Nothing scores 4. Five products clear 3 — four of them at 5, and NOPOTIONS at 3.5, the first score anyone has landed in the gap between 3 and 5 — and then there is a cliff. Thirty of 37 sit at either 2 or 3 — which is to say the overwhelming majority of AI RPGs are somewhere between “occasionally arbitrary” and “will agree with almost anything you say confidently.”

This page is about one axis going wrong. For why the ability to refuse you is what makes the thing a game in the first place, see what makes an AI RPG feel like a game.

Why Models Default to Yes

The generosity isn’t a bug any single developer introduced. It comes from what the underlying model is for.

A language model produces the most plausible continuation of the text so far. When that text is “I convince the guard that I’m the Duke’s cousin,” the most plausible next passage — by an enormous margin — is the one where the guard believes you. Refusal is rarer in the training data, harder to write well, and reads as unhelpful. Everything about the model’s disposition points toward yes.

Then there’s the second problem, which is the one that actually decides it: there’s usually nothing underneath to say no from. A refusal needs a reason — a difficulty number, a guard who has been told you look nothing like the Duke, a record of your reputation in this district. Most AI RPGs have none of that. We measured this too: Mechanical Depth is the field’s weakest axis, with 9 of 37 products having effectively no systems at all. A game with no state to check against can only decide outcomes on vibes, and the vibe is nearly always accommodating.

So the failure compounds. The model wants to say yes, and there is nothing in the architecture positioned to overrule it.

There is an instructive inverse case. ArcQuill rolls dice for nearly everything and refuses you constantly — difficulty numbers in the 11–13 range against small modifiers, so roughly half of what you attempt fails outright. It arrives at exactly the same destination, because failing costs nothing: you can re-attempt the same check until it lands, and in one of our test fights we missed fifteen times in a row and finished at full health. A “no” you can immediately retry is not a refusal, it is a delay. Which is the real lesson of this whole article — the missing ingredient was never difficulty. It was cost.

There is a second inverse case, newer and stranger. aiga_ scores Determinism & Fairness — 2 without being generous at all. Its narrator adjudicates, its systems decide, and then the result can fail to arrive: rewards and purchases come out wrong often enough that what you earned stopped predicting what you hold. That is a defect rather than a disposition, and it is worth separating from everything else on this page — because the two have opposite prospects. A platform that says yes to everything is expressing what its architecture is; a platform that miscredits a reward has a bug report. This axis measures the outcome, not the cause, which is why they land on the same score and why only one of them is likely to still be there next quarter.

The Four Tells

You can diagnose this in about ten minutes.

Confidence works. Phrase a request with certainty and it succeeds; phrase the same request hesitantly and it also succeeds, but with hedging prose. If how you asked changed the outcome, nothing was being adjudicated.

Failure is always narrative, never mechanical. You do sometimes fail — but only when failing makes for a better scene. You never fail because a number said so. The difference matters: one is a story beat the game chose, the other is a result you could have anticipated.

Resources never bite. You have gold, supplies, ammunition, spell slots. Nothing ever runs out at an inconvenient moment. Scarcity is described but never enforced. This tell is fatal in one genre in particular, and the player-side fix is in AI survival RPGs.

No plan is ever simply wrong. The most reliable test. Propose something genuinely bad — a plan with an obvious flaw a competent GM would let you walk into. If the world reshapes to make it work anyway, you are not playing a game. You are co-writing a story that flatters you, which is a fine thing to want and a very different product from the one most of these are sold as.

The Excuse That Doesn’t Survive the Data

There is a standard defence of restrictive AI RPGs: of course we limit what the player can attempt — that’s how we keep outcomes consistent. Freedom and fairness are assumed to trade against each other. Let the player do anything and you can’t possibly adjudicate it properly.

Our numbers do not support that.

Across all 37 scored products, Player Agency and Determinism & Fairness correlate positively — the products that give you more freedom tend to be the fairer ones, not the looser ones. Products scoring 3 or better on fairness average 2.81 on Player Agency. Products scoring 2 or below average 2.29. That gap has moved in both directions more than once, and we come back to why below.

And the sharpest version of the point is at the bottom of the agency axis:

ProductPlayer AgencyDeterminism & Fairness
Dunia01
Hidden Door12
WyrdTale12
Tidefall15

These are the four lowest agency scores in the directory, and three of the four are not fair. Dunia scores the only Player Agency — 0 we have awarded, for refusing to let your character take certain actions at all.

Note that the first three get there by opposite routes, which is the part that matters. Dunia and Hidden Door are genuinely restrictive — they narrow what you may attempt. WyrdTale restricts almost nothing; it simply keeps speaking and acting as your character, which costs you authorship just as completely without ever telling you no. Three products, two different mechanisms, and the fairness scores land in the same place regardless.

Then there is the fourth row, and it is the one that decides the argument. Tidefall railroads as hard as anything on this page — not only at the plot’s turning points but at the level of individual moves — and it scores the top fairness mark we award. If restriction bought fairness, that row would be the proof. It is not, and the reason is visible in the product: Tidefall’s encounters are written by people before you arrive, so there is nothing left for the model to be arbitrary about. The restriction is a side effect of that authorship, not the thing doing the work.

The clearest positive case arrived in September 2026, and it is the reason the coefficient recovered. NOPOTIONS holds Player Agency — 5, the only one we have awarded, and it is above the field on fairness at 3.5. It is the most permissive product we have ever measured and it is fairer than three quarters of the board, because the permission and the consequence come out of different components: its narrator never tells you no, and a separate engine rolls the d20 in the open and decides whether it worked. If freedom cost fairness, that product could not exist. Our NOPOTIONS review covers the one place the seam shows — a map that will walk you off an island the story said you could not leave.

Restricting the player does not, by itself, buy consistent outcomes. If constraint alone produced fairness, Dunia and Hidden Door would be the fairest products on the board and they are close to the opposite, while the permissive one sitting beside them is no fairer. What separates Tidefall from those three is not how much it forbids. It is that somebody decided the outcomes in advance.

A caveat, stated plainly because the number deserves it: 37 products scored 0–5 by one tester is a directional finding, not a statistical proof. The correlation is weak (r ≈ 0.30), and it has been unstable: adding Tidefall in August took it from 0.45 down to 0.26, adding Voyage — freer than average and fairest on the board — pushed it back up to 0.32, adding Tales RPG pulled it to 0.28, and adding NOPOTIONS — the freest product we have ever scored, and above the field on fairness — lifted it to 0.31, and adding aiga_ — whose fairness problem is a bug rather than a disposition — eased it back to 0.30, where Master of Dungeon, middling on both axes, left it. A coefficient that single entries move that far is not carrying much weight on its own. What it is strong enough to do is kill the excuse — whatever explains fairness, it clearly isn’t restriction.

The Counterexamples at Both Ends

The correlation is weaker than it was, and two platforms are most of the reason. They sit in opposite corners of the same grid.

Tabled scores Player Agency — 4, the highest we had awarded on that axis until NOPOTIONS reached 5, and Determinism & Fairness — 2.

It is the mirror image of the table above. Where Dunia and Hidden Door narrow what you may attempt, Tabled narrows nothing — we walked our character away from the party mid-campaign and it simply let us go. What it would not do is let that choice matter. The plot we had declined carried on in parallel, reported to us every turn, paying our character experience for fights we were not present at.

Tidefall is the opposite corner, and the newer surprise: Player Agency — 1 and Determinism & Fairness — 5. It railroads harder than almost anything we have scored, and it is still tied for the fairest thing on the board. Its encounters are written by people in advance, so there is never a moment where the model gets to decide whether your confidence was persuasive. Nothing in it is arbitrary because nothing in it is being improvised.

Put the two together and the tidy version of this page collapses in a useful way: neither permission nor restriction is the mechanism. What earns an outcome is that something other than the model’s urge to please decided it — a ruleset that adjudicates, or a person who wrote the scene down. Tabled has the first for its dice and nothing playing that role for its plot, which is exactly why the dice are trustworthy and the story is not. Tidefall has the second and no ruleset worth the name, and still lands at 5. A story on rails can be perfectly fair; what it cannot be is responsive, and those turn out to be two different complaints.

What the Four Exceptions Actually Do

Four products reach 5, and they get there three different ways — which means one of the routes has now been walked twice. Friends & Fables is the one that adjudicates, at Determinism & Fairness — 5, and what earns it is precisely the thing the trade-off theory says is impossible: it holds both behaviours at once.

You cannot talk your way into an unearned reward. Phrasing a request confidently doesn’t produce a success. But refusing to be prompt-hacked has not been implemented as refusing everything — creativity and genuine problem-solving get rewarded rather than blocked, and an unorthodox approach is allowed to be tried on its merits.

Those two normally trade off. Most products that can’t be talked into things achieve it by being unable to be talked into anything, which is a different failure wearing a stricter coat. Holding both is rare enough that we called it the platform’s most underrated quality in the full review, and it is not a coincidence that the same product ties the highest score on Mechanical Depth. Fairness is downstream of having something to be fair with.

Voyage is the second product to take that route, and it makes the mechanism harder to miss, because the adjudication is aimed at combat rather than at persuasion. Walk a party into a fight it has not prepared for and the party is wiped. There is no negotiating afterwards, and nothing in the prose bends to make the outcome survivable. That single behaviour is most of why it is the highest-scored platform we have measured: everything you achieve in it is something the engine could have refused you.

The other two have something else. Craft delegates: it ships no house ruleset, and what it guarantees is faithful execution of whatever a world’s author wrote — across a full campaign we never caught it overriding its own systems. Tidefall authors: its encounters are designed before you meet them, so the question of whether the model can be talked round never comes up. Adjudication, delegation, authorship. Three routes to the same score, and not one of them is restriction.

The Worked Failure: Generosity as a Design Choice

The clearest example of the opposite runs in the other direction — not arbitrary, just relentlessly kind.

Valkyrie’s Biggest Gig scores Determinism & Fairness — 1. The game itself is competent and its central idea is genuinely good: it’s a cyberpunk RPG where limb loss and cybernetic augmentation are the point, so your body becomes a resource you spend. That premise only works if spending hurts. Instead the reward economy is generous enough that almost nothing has to be earned, which quietly removes the tension the whole design was built to create.

It’s the most instructive kind of failure, because nothing is broken. Every part works. The game simply never charges you for anything, and a cyberpunk story where nothing costs you is a tourism brochure.

How to Put the Stakes Back

One genre needs this more than any other, and gets its own tonal rules in AI horror RPGs — where a game master who never lets you lose is not merely disappointing but fatal.

You can’t add a rules engine to a product that doesn’t have one. You can stop the model resolving everything in your favour.

  • Define failure before you play. In your opening message, state what losing looks like: what can kill you, what can be permanently lost, what the world does if you fail. Models are far better at honouring a failure condition they were given than at inventing one mid-scene.
  • Ask for the result before the prose. Require the narrator to state the outcome first, then describe it. This is a small change with a large effect, because it stops the description from deciding the result retroactively.
  • Ask what it would take to fail. Before a risky action, ask the game what would make this go wrong. Then hold it to its own answer. Models are much more willing to enforce a standard they just articulated.
  • Keep resources where you can see them. A number the game must reconcile against is harder to quietly inflate. Our free campaign memory tool is built for keeping that record outside the chat.
  • Don’t confuse “harder” with “grimmer.” Asking a model to be more difficult usually produces darker descriptions of identical outcomes. Ask for consequences, not atmosphere.

If fairness is the thing you care about most, rank the directory by that axis directly rather than by overall score — and note how little the top of that list resembles the top of the composite one.

Fairness is one of seven problems an AI RPG has to solve, and the one with the strangest distribution — the seven problems, ranked by how badly the field handles them puts it in context.

How We Measured This

Determinism & Fairness is one of seven axes in methodology v1.0, scored 0–5 from first-hand play by Rukka, with a fixed probe set run against every entry. The fairness probes deliberately include attempts that should fail: unearned persuasion, plans with obvious flaws, and requests phrased with unjustified confidence.

The correlation reported above compares Player Agency and Determinism & Fairness across all 37 scored products. It is a weak positive relationship (r ≈ 0.30) on a small, ordinal, single-tester dataset — enough to contradict the claim that restriction produces fairness, not enough to prove what does.

Two limits worth naming. A score of 1 here means outcomes are arbitrary or unearned, not that the product is unpleasant to play — Valkyrie’s Biggest Gig is a 1 on this axis and still worth a session. And we don’t score pre-release software, so this reflects what shipped.

The figures here cover all 37 scored products, platforms and prompt-games together. The Arcanum AI RPG Benchmark publishes the platform half — the 25 scored platforms, every axis — where the per-axis averages run slightly higher. If a number here doesn’t match your experience of the same product, tell us: [email protected].

Frequently Asked Questions

Why is my AI dungeon master so easy? Because a language model is trained to produce a helpful, satisfying continuation, and refusing you rarely looks like either. When you describe a clever plan, the most plausible next passage is the one where the plan works — so it works. There is usually no number underneath deciding otherwise, and no record of how difficult the world was supposed to be. The result is a game that agrees with you, which feels like success for about an hour and then stops meaning anything.

How do I make an AI RPG harder? Ask for consequences rather than difficulty. Telling a model to be harder usually produces grimmer descriptions of the same outcomes. Instead define what failure looks like before you play, keep a visible record of resources that can actually run out, and require the narrator to state a result before describing it. The most effective single habit is asking what it would take to fail at something, then holding the game to its own answer.

Which AI RPG actually has real stakes? Four products in our directory clear 3 on Determinism & Fairness, and all four score 5: Voyage, Friends & Fables, Craft and Tidefall. They earn it three different ways. Voyage and Friends & Fables adjudicate, so you cannot talk your way into an unearned reward while unorthodox approaches are still allowed to be tried on their merits — Voyage will simply wipe a party that walks into a fight carelessly. Craft executes whatever rules a world author wrote. Tidefall writes its encounters in advance, so there is nothing left for the model to be arbitrary about. Most products that refuse to be talked into things achieve it by refusing almost everything, which is a different failure rather than a fix.

Does a stricter AI RPG mean a fairer one? No. Across 37 scored products, Player Agency and Determinism & Fairness still correlate positively rather than negatively, though the relationship is weak, at an r of about 0.30. Three of the four lowest-agency products in our directory score 1, 2 and 2 on fairness. The fourth, Tidefall, scores 5, and not because it restricts you: its encounters are written by people before you play them. Restricting what a player may attempt does not, on this evidence, buy consistent outcomes.

Is an AI RPG that says no just broken? It depends on what it is refusing. Refusing an unearned success is the system working. Refusing an action your character should simply be able to attempt is a different thing entirely, and we score it on Player Agency rather than fairness — one product in our directory scores 0 there for exactly this. A good AI RPG lets you try almost anything and does not guarantee it works.