What Independent Research Says About AI RPGs — Where It Agrees With Our Benchmark, and Where It Doesn't
Independent academic research, testing AI models with methods nothing like ours, reaches the same conclusions our benchmark reaches on shipped platforms — on three of four major questions. On the fourth, how well AI keeps characters in character, the research is harsher than our scores. That disagreement turns out to be the most useful part.
This matters for a simple reason. Our benchmark has real limits, and we publish them: every score comes from one rater, and the test items are kept private so platforms can’t tune to them, which also means nobody else can re-run them. A single source, however careful, is still a single source. The strongest check on it is to find people who studied the same problems in a completely different way, without knowing we exist, and see whether they found the same thing.
They did. Here is where, where they didn’t, and what that tells players and developers.
Disclosures: several platforms in this dataset gave us free access or credits for review — Craft, Voyage, Tidefall, NOPOTIONS, Tabled, Tales RPG, aiga_, ArcQuill and AI Game Master. Fable Forge holds a paid sponsored placement on our directory and Questner bought an audit from Arcanum Forge, our design service; neither is scored here, because a rated platform may never be a client. None of it buys coverage or a score — how we handle this.
Two Kinds of Evidence
The research and our benchmark measure the same problems from opposite ends.
The research tests models. A typical paper takes several language models, puts each one in a controlled setup, and checks the output automatically or with a judge: a narrator against an automated player that tries to break the story, or a character against questions it should refuse to answer. That approach has real strengths we don’t have. It runs many trials, anyone can re-run it, and it changes one thing at a time.
Our benchmark tests platforms. We play the product a player actually buys — the engine, the memory system, the rules and the model together — and score it on the same seven axes every time. That has one strength the research doesn’t: it measures what is actually shipped, side by side. None of the papers in our research index compares released AI RPG platforms on one rubric.
When two methods this different land on the same answer, the answer is probably real. When they disagree, one of them is seeing something the other can’t. Both are useful.
Everything below was checked against each paper’s own text on 25 September 2026. The full list — 63 papers, each with what it measures, what it found and what it can’t see — is in the research index.
Agreement One: Long Campaigns Break
What our benchmark found. Across 27 platforms, two of the seven axes have never been topped: Memory & Continuity and Longevity. The best Memory score anywhere on the board is 4.25, and the best Longevity score is 4. The field averages 2.97 on Memory and 2.65 on Longevity. In our state of the field report we put it as “solved at turn one, unsolved at turn fifty”: the first hour is often excellent, and the problems only appear once a campaign is long enough to need its past.
What the research found. The same thing, measured precisely.
- NCP-Bench, accepted at ICML 2026, tests whether a narrator keeps its story’s facts and commitments while an automated player pushes against them. The best model tested kept its story intact in only 42% of runs after 20 turns, and fact-conflict rates ran from 40% to 68% across models.
- RoleBreak tested voice role-play over long conversations. Even the strongest system broke character after 10.4 turns on average. Bigger models delayed the failure; none removed it.
- ConStory-Bench, which studies contradictions in long stories, found errors cluster in factual and temporal details and tend to appear around the middle of a narrative — after the setup, before anything resolves. That is precisely the stretch of a campaign our Longevity axis probes.
Why this is convergence rather than coincidence. The research uses automated adversaries, short fixed horizons and bare models. We use human play, long campaigns and complete platforms with their own memory systems. The research breaks faster because its setups are built to break things — 20 turns under attack is harsher than 50 turns of ordinary play. But both methods find the same wall, and neither has found anything that clears it.
Agreement Two: Good Writing Is Not a Working Game
What our benchmark found. The category’s strongest axes are the ones about writing: NPC Fidelity averages 3.20 and Signature Design 3.26. Its weakest is Mechanical Depth, at 2.52 — whether there are real systems under the prose. The field, as our state-of-the-field report puts it, “solved characters and never solved consequences.”
What the research found.
- RPGBench asks models to build a text RPG and then run it, checking every state update automatically. Its conclusion, in its own words: models “can produce engaging stories but often struggle to implement consistent, verifiable game mechanics, particularly in long or complex scenarios.”
- NCP-Bench adds the sharper version: high writing quality did not predict whether a model kept the story consistent. A narrator can be fluent and wrong at the same time.
This is the finding players feel most and name least. A session that reads beautifully can still lose your inventory, forget a wound or let you win a fight you should have lost. The prose and the game are separate achievements, and the research shows they come apart even inside a single model.
Agreement Three: Game State Outside the Model Helps
What our benchmark found. The only two platforms above 4.0 are Voyage (4.4) and Craft (4.3). Both are built around state the narrator does not own. Voyage’s World Engine tracks skill checks, hit points and positioning outside the prose. In Craft, a worldbuilder can code new gameplay systems rather than describe them.
The wider board shows the same pattern. Our Mechanical Depth axis measures whether real systems run under the prose. The four platforms scoring 4 or higher on it average 3.72 overall. The fourteen scoring 2 or lower — mostly pure narration — average 2.46.
Two cautions keep that honest. Mechanical Depth is one of the seven axes that make up the overall score, so part of that gap is the axis counting itself. And the pattern has an exception: AI Realm runs real systems and still scores 2.3 overall, because systems alone don’t make a good game. It is a pattern in 27 data points, not proof. The research supplies what our data can’t: the mechanism.
What the research found. Paper after paper improves consistency the same way — by taking the world out of the model’s hands.
- FIREBALL collected about 25,000 real Dungeons & Dragons sessions with the true game state recorded, and showed that giving models that state improved the turns they generated.
- A study of AI game masters using function calling — the model calls game functions instead of narrating outcomes — improved both narrative quality and state consistency, by human evaluation and by unit tests.
- PAYADOR and IVIE ground the model in a symbolic world state, and found that this keeps worlds consistent without removing creative freedom.
- Orchestrated Reality, a 2026 work in progress, states the problem as bluntly as anyone has: in deployed systems, the narrator “asserts state in free prose without any validated representation.”
The research explains why the pattern in our data exists, and our data shows the idea working in products people can play today. Neither is conclusive alone. Together, they make the strongest case in this article.
The Disagreement: Characters
Here the research and our benchmark part ways, and it’s worth being precise about how.
What our benchmark found. NPC Fidelity is one of the category’s strongest axes. Three platforms hold a 5 on it. Our reading has been that the category largely solved characters before it solved games.
What the research found. Characters are much more fragile than our scores suggest.
- RoleBreak: first persona failure after about 10 turns, in the strongest system tested.
- Too Good to be Bad: models lose fidelity steadily as a character becomes less moral. Safety-trained models swap real malice for surface aggression, and general chatbot skill does not predict villain skill.
- CHARM: characters often recognise that a question is outside what they could know, and answer it anyway.
- ArcANE: characters follow their own development best when the system hands them their arc explicitly. In every model tested, that beat giving them the relevant past events alone.
Why the two disagree. We think three things explain most of it.
- The research attacks; we play. Much of this work uses adversarial prompts designed to push a character out of role. We play as a player would. A character that holds up in ordinary play can still break the moment someone deliberately leans on it. Both are true.
- Platforms add scaffolding the research strips away. Our top NPC scores go to products with character cards, memory systems and prompt design around the model. Most papers test the model with a short persona description. The difference between those two setups is exactly what a platform is selling.
- Our own work already found part of this. Our NPC design analysis documents three failures that appear once a second character enters a scene, including one we called “the character becomes your mirror” — a character that slides toward agreeing with you. “Too Good to be Bad” is the model-level version of that finding: the pull toward pleasing the user is strongest where a character is supposed to resist.
What we take from it. Our NPC Fidelity scores describe characters in normal play, and they should be read that way. The research is a fair warning that “solved” was too strong a word: characters hold under ordinary conditions and fail under pressure. An adversarial probe of that kind is something our method does not currently include. Any change to how we score would be announced on the methodology page first, never applied quietly.
What the Research Sees That We Can’t
Some findings sit outside anything our benchmark measures, and they matter anyway.
AI characters are not automatically more fun. In a randomised study of 130 players, NPCs driven by a language model significantly raised cognitive load and did not significantly improve the overall experience compared with scripted NPCs. Players felt more autonomy, but rated usability and trust lower, and the extra effort was largest in open-ended tasks like relationship building. This is one prototype game, not a verdict on the category. But it is a useful correction to the assumption, common in marketing, that a talking NPC is automatically a better NPC. The freedom has a cost, and the game has to be designed to carry it.
AI characters can be tricked, and can be biased. One study shows that prompt injection can make an NPC reveal secrets it was meant to keep. That stops being theoretical once a game ships characters who chat live, as Where Winds Meet now does — and its players found a version of the trick within days. FairGamer found social bias in the decisions of all seven frontier models it tested, with larger models showing more of it.
Latency is real. A VR study of GPT-4-driven characters measured an average response cycle of seven seconds, growing as the conversation got longer. Anyone designing for voice needs that number.
What This Does and Doesn’t Prove
It is worth being careful here, because overclaiming would undo the point of the exercise.
It does show that the problems our benchmark measures are real problems, found independently by people using entirely different methods, and that our three central findings about the category agree with the research.
It does not validate any individual score. The research never tested the platforms we rate. Agreement about the field says nothing about whether a particular platform deserves a 3.5 or a 4.
It does not remove our limits. Our scores still come from one rater, on test items nobody else can re-run. The research is a check on our conclusions, not a second rater. That is why we publish the rubric, the data and a permanent snapshot of every edition — so the conclusions can be checked even where the scores can’t be repeated.
Some of the research is early. Several papers are preprints, some studies are small, and one we cite is explicitly a work in progress. The research index shows which is which, and states a conference only when the authors themselves do.
What It Means for Players
- Expect long campaigns to drift, whatever you play. No platform and no model has solved this. The practical defences are the same everywhere: keep important facts in the platform’s permanent notes rather than trusting the story to remember them. Our guide to why AI campaigns fall apart covers the fixes.
- If rules matter to you, pick a platform with tracked state. Both our data and the research say the same thing: a world the narrator can’t quietly rewrite is more consistent.
- Villains and antagonists are a known weak spot. If a character you expected to resist you keeps agreeing with you, it isn’t just your platform. It is a documented limit of safety-trained models.
What It Means for Developers
- Hold the world outside the model. This is the most consistent result in the research and the clearest pattern in our data.
- Test long and test adversarially. A game that is excellent for 10 turns tells you very little. The research’s failure points — 10 turns under pressure, 20 turns of adversarial play — are a better bar than a good first session.
- Budget for cognitive load. Open-ended conversation is effort for players. The 130-player study suggests the design around the AI matters as much as the AI.
- The research is free. Every paper in our index links to its source, and the whole index is downloadable as JSON.
How We Did This
We searched arXiv for research on role-playing, game masters, text adventures, interactive fiction and AI-driven characters, read about 255 abstracts, and kept 63 papers that bear on at least one of our seven axes. Titles, authors and dates come from the arXiv record itself. We state a venue only when the paper’s authors state it: one well-known paper says it was submitted to a major conference, and a submission is not an acceptance. A script re-checks every entry against its source before each update. The full list, with what each paper measures and what it can’t see, is the research index. Our own measurements are on the benchmark.
Frequently Asked Questions
Does academic research agree with the Arcanum benchmark? On three of four major findings, yes. Research on AI models and our first-hand measurement of shipped platforms both find that long campaigns break down, that good writing does not mean a working game, and that game state tracked outside the model improves consistency. On character fidelity the research is harsher than our scores, and we explain why in the article.
What research exists on AI RPGs and AI game masters? The closest work includes RPGBench, which tests language models as text RPG engines; NCP-Bench, which tests whether a narrator keeps its story commitments over long play; FIREBALL and related Dungeons and Dragons research on game state; and many benchmarks for role-playing characters. Our research index lists 63 papers, each mapped to the benchmark axis it bears on.
How long can an AI keep a story consistent? Not long, under pressure. In NCP-Bench the best model kept its story intact in only 42 percent of runs after 20 turns, and in RoleBreak the strongest spoken system broke character after about 10 turns on average. Our own benchmark finds that no shipped platform reaches a top score on Memory and Continuity or on Longevity.
Do AI-driven NPCs make games more fun? Not automatically. In a randomised study of 130 players, NPCs driven by a language model raised players’ cognitive load and did not significantly improve their overall experience. Players felt more autonomy, but rated usability and trust lower. The effect depends heavily on how the game is designed around the AI.
Does this prove the Arcanum benchmark is right? No. Agreement between different methods is strong evidence that the findings are real, but it does not validate any single score. Our scores still come from one rater, and the research tests models rather than the platforms we score. What it shows is that the problems we measure are the same ones independent researchers find, by different routes.