Why Our AI RPG Ratings Changed: We Re-Scored Every Platform, and 13 of 15 Went Down

In July 2026 we re-tested every AI RPG platform in our directory against a published seven-axis rubric. Fifteen platforms were re-scored. Thirteen went down, two went up, and not one stayed the same.

That is an uncomfortable thing to publish about your own back catalogue, so it is worth being precise about what it does and does not mean. It does not mean the category got worse over three weeks. It means we replaced a number that was one person’s overall impression with a number that has seven separate judgements underneath it — and the two instruments disagree, systematically and in a direction that flatters nobody, least of all our earlier selves.

The full board is now published as the Arcanum AI RPG Benchmark. This is the note explaining how it got there.

The Old Numbers Were One Person’s Impression

Until this month, an Arcanum rating worked the way almost every roleplay rating on the internet works. Rukka played a platform, formed a view, and wrote down a number between 1 and 5. Nothing was fabricated and nothing was second-hand — the first-hand rule has been in place since the site launched — but the number itself was holistic. It averaged everything he had noticed, weighted by whatever had stuck, and then reported the average as though it were a measurement.

The problem with a holistic number is not that it is dishonest. It is that it is unfalsifiable. If you had written to us in June and said our Dunia rating was too high, there would have been no working to check, because there was no working. There was a verdict.

The methodology we published this month exists to fix exactly that. Seven axes, each scored 0–5 on its own evidence, each answering a question narrow enough to argue about:

  • Memory & Continuity — does the world remember what happened, and for how long?
  • Player Agency — can you act off-path, and does the narrator stay out of your character’s mouth?
  • NPC Fidelity — consistent personality, own goals, durable reactions, or agreeable mirrors?
  • Mechanical Depth — real systems underneath, or narration wearing dice?
  • Determinism & Fairness — outcomes earned and consistent, or arbitrary and luck-washed?
  • Longevity — does it survive past turn 50 and the novelty window?
  • Signature Design — does it do something nobody else does?

The headline rating is now the mean of the axes that apply. It is computed, not chosen. When an axis moves, the rating moves with it, and a score can no longer drift away from the reasons given for it.

What Changed Is the Instrument, Not the Products

Here is the whole re-score, sorted by how far each platform moved.

PlatformWasNowChange
Dunia4.01.7−2.3
Perchance3.01.4−1.6
Deep Realms4.02.7−1.3
AI Realm3.52.3−1.2
FableAI3.52.4−1.1
MacerAI4.02.9−1.1
Janitor AI4.03.0−1.0
Old Greg’s Tavern4.03.0−1.0
AI Dungeon3.52.6−0.9
Infinity DM3.52.7−0.8
Friends & Fables4.53.9−0.6
Hidden Door3.02.7−0.3
Chub AI2.01.9−0.1
Character.AI2.02.1+0.1
AI Game Master1.02.3+1.3

Exactly one of those fifteen moved because the software changed. It is the one at the bottom, and it went up.

The Biggest Drop Was Ours to Take

Dunia fell further than anything else on the board, and the platform was not touched between the two reviews. Same product, same build, same everything. What changed is that the new method includes a probe set — a fixed sequence of things we try on every platform — and free play had never happened to try one of them.

The agency probe found that Dunia refuses to let your character take certain actions at all. Not “handles them awkwardly.” Refuses. That is the axis scoring 0, the only zero anywhere in the directory, and it is the difference between a 4.0 and a 1.7 once six other axes are averaged in beside it.

We had played Dunia for hours and never hit it, because when a game gently steers you, you tend to go where it steers. The probe does not go where it is steered. That is the entire reason it exists, and this is the clearest evidence we have that the method earns its keep: a rubric overturned the reviewer’s own published verdict on evidence he would never have collected otherwise.

Everything the original Dunia review praised is still true and still praised. The prose is good. The branching editor is genuinely clever. The review now says all of that and the refusal too, which is a more useful page than the one that only said the first half.

The One That Went Up

AI Game Master was, in July, a 1 out of 5 on this site, under a review titled “Why Not Just Play AI Dungeon?” It is now 2.3, and that change is real rather than instrumental: the platform was substantially rebuilt between the two visits, and it now ships items, a quest log, party members and a chapter economy that the version we originally reviewed simply did not have.

The developers provided free access for the re-review. That review is independent and unpaid, and no commercial relationship exists between AI Game Master and Arcanum.

We are keeping this one visible rather than quietly swapping the number, because a platform going from bottom-of-the-directory to competent is a better story than either score alone — and because it is the only real test of whether a review site tracks products or just freezes a verdict and moves on. Ours froze one for three weeks. The re-score unfroze it.

What the Old Scale Was Actually Hiding

The interesting finding is not that the numbers fell. It is what they fell into.

Five platforms carried an identical 4.0 under the old system: Dunia, Deep Realms, MacerAI, Janitor AI and Old Greg’s Tavern. Under seven axes those five now read 1.7, 2.7, 2.9, 3.0 and 3.0 — a spread of 1.3 points between products the old scale said were the same.

Two platforms shared a 3.0: Perchance and Hidden Door. They are now 1.4 and 2.7, which is the widest split of the lot.

So the old ratings were not merely generous. They were compressed. A 1-to-5 scale used holistically collapses toward the middle-upper band, because a platform that is excellent at one thing and poor at another averages out in the reviewer’s head before it ever reaches the page. Seven axes force the excellence and the poverty to be written down separately, where they stay visible. That is why Janitor AI scores 3.0 overall and holds the only 5 we have awarded on NPC Fidelity, and why AI Dungeon scores 2.6 while winning Player Agency outright. A single number destroyed both facts. The axes recover them.

There Is No Correction Factor

An obvious question: can you take an old Arcanum rating and adjust it to a new one?

No. The changes run from +1.3 to −2.3, and platforms that shared an old rating landed as much as 1.3 points apart. There is no offset that reproduces that, because the old numbers did not encode the distinction in the first place — you cannot recover information that was never written down. Every entry had to be played and scored individually, which is why it took a month — and why the twelve prompt-games and Custom GPTs in our games directory went through exactly the same process rather than being converted on paper.

This also means old and new numbers are not comparable to each other, and we do not compare them anywhere except on this page and in the dated note at the top of each re-scored review. If you find a page on this site putting a legacy number beside a v1.0 number as though they measured the same thing, that is a bug — tell us.

What This Means When You Read a Score Here

Three practical things.

The composite is the least interesting number on the page. It averages seven judgements and then tells you none of them. Read the axis row instead — the compare table sorts the whole directory by any single axis, which is usually the question you actually arrived with.

A lower score is not a worse platform than it was in June. It is a more precisely described one. Nothing on this site got downgraded for cause except where the review says so explicitly.

The scores are falsifiable now, and we would like them falsified. Every axis has a published question behind it and a paragraph of reasoning on the entry. If a number does not survive contact with your experience of the same product, that is worth an email — disagreeing with a verdict is not a correction, but a score whose reasoning does not hold up very much is.

The full board, including the three axes on which no platform has yet scored top marks, is at the Arcanum AI RPG Benchmark. The rubric, the tiers and the places the method is weakest are on the methodology page.

Frequently Asked Questions

Why did Arcanum’s AI RPG ratings go down? Because we replaced the way we measure, not because the platforms got worse. Every old rating was a single number set by overall impression after free play. Every new rating is the mean of seven separately scored axes — memory, player agency, NPC fidelity, mechanical depth, determinism and fairness, longevity, and signature design — measured with a fixed probe set that deliberately pushes at things ordinary play never attempts. Thirteen of fifteen platforms scored lower under the new method, two scored higher, and none scored the same.

Does a lower score mean the platform got worse? Almost never. Only one platform in the re-score changed because the product changed, and that one went up rather than down. The other fourteen moved because the instrument changed. Dunia is the clearest case: it dropped further than anything else on the board and the software was not touched between the two reviews. The probe set simply tested something free play never had.

Can you convert an old Arcanum rating into a new one? No, and we would not publish the conversion if you could. The changes run from plus 1.3 to minus 2.3, so there is no offset that would work. Platforms that shared an identical old rating landed as far as 1.3 points apart under the new method. That spread is the entire point: the old scale was compressing genuinely different products onto the same number, and no correction factor can recover a distinction that was never recorded.

Which platform’s score changed the most? Dunia fell the furthest, from 4.0 to 1.7. AI Game Master rose the furthest, from 1.0 to 2.3, and it is the one case where the product itself had genuinely been rebuilt between reviews. Perchance fell the second furthest, from 3.0 to 1.4, which makes it the lowest-scoring entry in the directory.

Where can I see all the new scores? The Arcanum AI RPG Benchmark publishes the full board — every scored platform, every axis, every score, sorted by composite. It has grown since this re-score and keeps growing; each platform enters the board the moment it earns a rating. The methodology page publishes the rubric itself, the scale, the test tiers and the places where the method is weakest. Individual review posts carry a dated note explaining what a platform used to score and why the number moved.