The Next AI RPG War Won't Be About Better AI
For about six years, the winner of any AI roleplay contest was whoever had access to the better model. That era is ending, and the scored evidence says it is ending faster than the marketing has noticed.
The bet was rational while it lasted. In 2019 the difference between a good AI text adventure and an unplayable one really was the model, because the models were bad enough that the gap between them swamped everything else. Prose quality was the product. Whoever got the better weights won.
Look at what we have measured since, and that stopped being true somewhere in the last year. Across 37 products scored on seven axes from first-hand play, which model a platform runs is a poor predictor of how good the platform is. What separates them now is the engineering around the model, the ideas it is asked to serve, and the discipline to stop it doing what it wants to do.
Disclosures: several platforms named here gave us free access or credits for review — Craft, Voyage, Tidefall, NOPOTIONS, Tabled, Tales RPG, aiga_, Master of Dungeon, ArcQuill and AI Game Master. None of it buys coverage or a score — how we handle this.
The Cleanest Proof Is at the Two Ends of Our Board
Character.AI trains its own proprietary model. It is one of the largest companies in this category, and the model is genuinely good at the thing it was built for — it holds NPC Fidelity — 4, better than most of the board. Its composite is 2.1 out of 5.
Craft rents. Its model band is published in its own patch notes and revised there — MiMo 2.5, DeepSeek v4 Flash, Hy3 at launch, GLM 5.3 Flash since. Commodity models, available to anyone with a card. Its composite is 4.3, second on our board, and it holds Mechanical Depth — 5 and Signature Design — 5.
The company that built its own model scores 2.1. The company renting models you could rent this afternoon scores 4.3.
That is not an argument that training your own is a mistake — Voyage runs Latitude’s own model and sits first at 4.4. It is an argument that the model is not where the difference is being made. Both ends of our board contain both approaches.
The Controlled Experiment Nobody Designed
Our prompt-games are an accident that turned into a natural experiment. They are text files — a prompt you paste into a subscription you already pay for. Every one of them runs on the same commercial models anyone can use, with no engine, no backend, and nothing between the author’s ideas and the model.
If the model were the product, they would all score the same. They do not.
| Model they run on | Games | Range |
|---|---|---|
| ChatGPT | 6 | 2.0 – 2.6 |
| Claude | 3 | 1.6 – 2.3 |
| Gemini | 3 | 1.9 – 2.1 |
The spread within a single model is as wide as the spread between models. And on individual axes it is much wider: the three Claude prompt-games span Signature Design 0 to 3 and Mechanical Depth 0 to 3 — most of the usable scale — on identical capability, identical access, identical everything except what the author wrote.
The only variable is the person who wrote the file. That is the whole thesis in one table.
The Field Has Quietly Stopped Advertising Its Model
Here is a softer signal that turns out to be a loud one.
Nine of our 25 rated platforms do not publish which model they use at all. Not vaguely — explicitly. Their entries in our directory read “Not named,” “Undisclosed,” “Not published,” because we went looking and there was nothing to find.
| Platform | Score |
|---|---|
| NOPOTIONS | 3.8 |
| Master of Dungeon | 3.3 |
| Tidefall | 3.3 |
| Tabled | 3.1 |
| Old Greg’s Tavern | 3.0 |
| aiga_ | 2.9 |
| Tales RPG | 2.8 |
| FableAI | 2.4 |
| Dunia | 1.7 |
Two things stand out. They span almost the entire quality range, from the fourth-best platform we have scored to nearly the worst — so hiding the model is not a tell for either good or bad. And more importantly: in a market where the model was the product, nobody would hide it. Every other technology category in the last decade has been drenched in powered by badges. This one increasingly is not, because the vendors know the model is not the thing they are selling.
The exceptions prove it rather than break it. ArcQuill lets the player pick from five models and publishes a reliability score for each — that is a design decision about transparency, not a capability claim. AI Dungeon publishes a large roster because model choice is its product surface. Neither is saying “our model is better.”
The Ceiling That Fell to Craft, Not Capability
The strongest evidence we have is a prediction of ours that broke in public.
In July 2026 we deposited our benchmark dataset on Zenodo with a permanent DOI, and its README made a falsifiable claim: three of the seven axes — Memory, Longevity and Player Agency — had no top score anywhere in the field, and they were unclaimed because they needed the underlying technology to improve. The claimed axes were the ones a competent team could solve with craft. The unclaimed ones were waiting on better models.
Two months later NOPOTIONS scored Player Agency — 5, the first perfect score ever awarded on that axis, from a small independent studio.
It did not get there with a better model. It got there with restraint — a narrator that describes a scene and then stops talking, on a map open enough that there is nowhere to railroad you back onto. That is a design decision and a piece of engineering discipline. The axis was not waiting on capability at all; we had simply misfiled it.
We cannot edit the deposited file, which is the entire reason for depositing it — the full account is here. But the lesson generalises further than the one axis. We took a problem we had classified as “needs better AI,” and a small team solved it with an idea. Having been wrong about that once, we now expect more of what looks capability-bound to be craft-bound.
The Honest Counterargument
Two axes have still never been maxed by anyone: Memory & Continuity, ceiling 4.25, and Longevity, ceiling 4. Our own benchmark still argues these need the technology to improve, and that argument is not obviously wrong — a language model genuinely has nowhere to keep a number between turns except the conversation itself, and that is a property of the architecture rather than a failure of effort.
So the strong form of the claim — models no longer matter at all — is not what the data supports.
The defensible form is narrower and still consequential. Within the range of models that actually ship today, the choice between them has stopped being the thing that separates products. A better model raises the floor under everybody at once, which by definition differentiates nobody. The differences that remain are in what teams build on that shared floor.
And there is a reason to think even the two remaining axes are more tractable than they look. The products nearest the top on memory got there by five different architectures — Voyage, Craft, Hidden Door, Tabled and WyrdTale — not by using a better model than each other. Five different engineering answers to the same problem is what a craft problem looks like.
What the New War Gets Fought With
If the model is a shared floor, the fighting moves to three places. All three are visible in the scores already.
Engineering: somewhere to keep what happened. Three of the seven problems — real mechanics, NPCs with goals that advance while you are elsewhere, and quests that survive you going off-path — are the same requirement wearing different clothes: a place to hold state that is not the conversation. It is the field’s worst axis, with Mechanical Depth averaging 2.27 and 24 of 37 products at 2 or below. It is also demonstrably solvable, because two products have reached full marks on it by opposite routes — Friends & Fables implementing a real ruleset properly, Craft letting world authors write systems in code. Why AI RPGs cannot do mechanics is the architecture argument in full.
Ideas: something to do that nobody else is doing. Signature Design is the axis with the most perfect scores in the field — five products have maxed it — and it is the one that correlates most strongly with a product’s overall quality of all seven. It is also the cheapest to attempt and the hardest to fake. It does not require a research lab. It requires knowing what your game is for.
Restraint: the discipline to build a system that can refuse the player. This is the least discussed and probably the most valuable. A model left to itself will agree with you, because agreeing produces the better next paragraph. Building something that says no — dice you cannot argue with, a rule the narrator does not get to reinterpret, an NPC who will not go along with it — is entirely a design problem. Nineteen of our 37 products sit at exactly 2 on Determinism & Fairness, which means the overwhelming majority of this field has not attempted it. We wrote about why that line is the one that makes it a game.
What This Changes for a Player
The practical version, because this is not only an industry argument.
Stop shopping for the model. “Which AI RPG has the best AI” is close to the wrong question now. It was a good question in 2022 and it has quietly become a bad one, and you can check that yourself: eight of the platforms we have scored do not tell you their model, and knowing it would not have predicted the score.
One clarification, because we publish the other side of this too. If you are choosing a raw model to roleplay with directly — in a chat window, through a prompt, in SillyTavern — then the model absolutely is the product, and which one you pick genuinely matters. Those are different questions. When you buy a platform, you are buying everything built around the model. When you paste a prompt into a chat window, you are buying the model and your own writing.
What to look at instead: what happens when you try to do something the game should refuse. That is a ten-minute test and it tells you more than any model roster.
What This Changes for a Developer
Three uncomfortable implications, stated plainly.
- A model upgrade is not a roadmap. It raises your floor and your competitors’ at the same moment. Anything that arrives for free arrives for everyone.
- Training your own buys control, not quality. Both ends of our board contain teams that did it. It is an enormous commitment that does not, by itself, appear in the scores.
- The biggest gaps left are ones a small team can close. The largest genuinely unclaimed space we have measured is an NPC with a goal that advances while the player is somewhere else — nobody has built it, and it does not need a frontier lab. It needs state and a scheduler.
The best-scoring things in this category are not the ones with the best AI. They are the ones where somebody decided what the game was for, built something to enforce it, and then had the discipline to let the model do only the part it is actually good at.
How We Measured This
The seven axes, the scale, the test tiers and the evidence standard are on our methodology page. Every score comes from playing the product first-hand, by Rukka, against a fixed probe set. The model information comes from each platform’s own site, docs or patch notes at the date on its directory entry, and where a platform publishes nothing we record that rather than guessing.
Three caveats worth stating rather than burying.
This is a snapshot argument, not a trend line. We can show that model choice does not predict quality today, across this field. We have two editions of this dataset, July and September, which is not enough to prove a direction of travel — the claim that models will matter less over time is a forecast, and it is ours, not a measurement.
Grouping platforms by “own model” versus “commodity model” is our reading of their disclosures, and several of those disclosures are thin. The two-ended comparison we lead with — Character.AI against Craft — is exact and checkable. Any group average would carry our classification inside it, which is why we have not published one.
We do not score models, only products. A platform’s score reflects its engine, rules, memory, interface and model together, deliberately, because that bundle is what a player actually buys. Nothing here should be read as a ranking of the underlying models against each other.
Frequently Asked Questions
Does the AI model matter in an AI RPG?
Less than almost anyone expects, once you are comparing finished products rather than raw models. Across the platforms we have scored, which model a platform runs is a poor predictor of how it scores. The clearest illustration is at the two ends: Character.AI trains its own proprietary model and scores 2.1 out of 5, while Craft runs commodity models anyone can rent and scores 4.3. The model sets a floor on prose quality. Everything above that floor is built.
Which AI RPG has the best AI?
It is close to the wrong question now, because the gap between the good commercial models has narrowed to the point where the surrounding engineering matters more than the choice between them. Nine of the 25 platforms we have scored do not publish which model they use at all, and they span almost the entire quality range. If you want the best AI RPG, look at what was built around the model rather than at the model.
Will better AI models fix AI RPGs?
Partly, and less than the field assumes. In July 2026 we published a prediction that three axes were unreachable until the underlying technology improved. One of them fell two months later, and not to a better model — NOPOTIONS reached a perfect Player Agency score through restraint and map design. That was craft solving something we had filed as a capability problem, and it is the reason we now expect more of the remaining problems to be craft than we did.
Why do AI RPGs running the same model score so differently?
Because the model is roughly a tenth of the product. Our prompt-games make this visible because they are the controlled experiment: they are text files running on the same commercial models anyone can use. The six on ChatGPT range from 2.0 to 2.6, and the three on Claude range from 1.6 to 2.3 with Signature Design spanning 0 to 3 on identical capability. Same model, same access, results across most of the usable scale.
What will AI RPG developers compete on in future?
State, adjudication and restraint — which are engineering and design problems rather than model problems. Somewhere to keep what happened that is not the conversation, something that decides outcomes and cannot be argued with, and the discipline to let a narrator stop talking. None of those arrive with a model upgrade, and all three are places where a small team has beaten a large one in our scoring.
Is it better for an AI RPG to build its own model?
On current evidence it is not the differentiator it looks like. The top of our board includes both approaches — Voyage runs Latitude’s own model at 4.4 and Craft runs rented commodity models at 4.3 — and so does the bottom. Training your own is a large, expensive commitment that buys control rather than quality, and quality in this category is coming from the engine, the rules and the restraint around the model.