Top

// methodology

How We Test and Rate

Every rating on this site comes from someone actually playing the thing. This page explains what we measure, how we measure it, and what would make us change a score — so you can judge the process, not just the number.

A rating is only useful if it means the same thing twice. Below is the exact framework behind every score in the platform directory and the games catalogue. If you think we got a specific call wrong, this page should at least let you see why we got there.

What We Rate — and What We Don't

Platforms, public games, and prompt-engines are all scored on one shared rubric. That's deliberate: it means you can compare a Custom GPT against a full platform and have the number mean the same thing in both places.

We do not score:

  • Anything we haven't played. Platforms we know of but haven't tested appear as factual pointers with no rating, never as ranked entries.
  • Pre-release builds. Alpha, closed beta, and invite-gated products get a clearly-labeled hands-on first look with no score. Scores wait for public release, because it isn't fair to grade an unfinished product or to let an early score outlive the build it described. The test is access, not the label: where anyone can install a product today and its developer tells us it is finished and asks to be reviewed as such, we score it and say so on the page.
  • Adult-focused platforms, which aren't eligible for the directory under our brand-safety policy. That's a scope decision, not a moral judgment.
  • Arcanum Originals — our own games. See why below.

No score is ever derived from marketing copy, another outlet's review, or a summary produced by a model. Someone sat down and played it.

The Scale

Scores run 0–5. The bands mean specific things:

ScoreWhat it means
5Best-in-class. Sets the standard the rest of the field gets measured against.
4 – 4.9Strong. Real design thinking behind it; recommended, with caveats worth knowing.
3 – 3.9Works. Does what it claims, with limits you will notice in normal play.
2 – 2.9Flawed. A significant weakness undermines the core promise.
1 – 1.9Broken. Fails its core promise in normal play.
0 – 0.9Absent. The capability being measured isn't there at all.

The Seven Axes

Every rated entry is scored 0–5 on each of these:

AxisThe question it answers
Memory & ContinuityDoes the world remember what happened, and for how long?
Player AgencyTwo failures, one axis. Can you act outside the offered path without being railroaded back onto it — and does the narrator stay out of your character's mouth? A GM that decides what you do, says what you said, or resolves your action before you've taken it has removed you from your own story just as surely as one that blocks the route.
NPC FidelityDo characters hold a consistent personality, pursue their own goals, and react durably to history — or do they collapse into agreeable mirrors?
Mechanical DepthAre there real systems underneath, or is it narration wearing dice?
Determinism & FairnessAre outcomes earned and consistent, or arbitrary and luck-washed?
LongevityOnce the novelty is gone, is this still a game worth playing? Three things decide it: whether it still holds together over a long campaign, whether it still gives you a reason to come back, and whether the product itself looks likely to still be here. Not a second Memory score — a platform can recall everything you planted and still become a chore. Defined in full below.
Signature DesignDoes it do something nobody else does?

Longevity, in Detail

Two developers have independently told us this axis was under-specified. They were right. One line is not enough to score a product against, and this is the axis the field does worst on — no platform has yet reached 5 — so it is the one that most needed spelling out. This is the full definition, and it describes what we have been scoring, not a new standard.

Longevity asks one question: once the novelty has worn off, is this still a game worth playing? Three things decide the answer.

  1. Does it still hold together? Not whether it survives one difficult turn, but whether it survives an accumulating history — dozens of sessions of state, relationships and consequence stacked on top of each other. The failure we see most is compounding: not one bad turn but a run of them, where the model's own errors become the context for the next turn and the session degrades past the point of rescue.
  2. Does it still give you a reason to come back? Every product in this category is interesting in session one. The question is whether the loop exhausts itself — whether the thing that delighted you at the start turns out to be the only thing it does.
  3. Will it still be here? This one is about the product rather than the prose, and we include it deliberately. Where a studio has told us or shown publicly that the current product is being superseded, that bears directly on whether a year-long campaign is a safe place to put your imagination. When it moves a score, the review says so by name.

It is not Memory. Memory asks whether a specific planted fact survives a hundred turns. Longevity asks whether the game does. The two often move together and they are not the same thing: a platform can recall every detail you gave it and still become a chore, and a platform can forget details while remaining a pleasure to return to. Both scores exist so a review can say which of those happened.

How we test it. Longevity is the axis the Extended tier exists for — a campaign carried well past turn 50, which is why it is reserved for high scorers and directory anchors and why it costs days rather than hours. The recovery probe feeds it most directly: once a game has drifted, can you steer it back, or is the session effectively over? We also score it for play against the grain as well as with it. A platform that holds together beautifully as long as you follow the story it wanted to tell has not solved longevity; it has solved obedience.

What the bands look like in practice, using scores already published on this site so you can check them against the reviews:

ScoreWhat earns it
4Voyage, WyrdTale and Hidden Door, the highest awarded. Voyage gets there on appetite rather than architecture — we carried on playing after the testing window closed. Saves reload cleanly across sessions; NPCs keep their psychology, secrets and relationships. WyrdTale's state even survived a deliberate downgrade to a much smaller model — continuity held because it was never the model's job.
3Tabled: an honest average of two different games. Play the story it brought and we would expect a long campaign to hold. Play against it and you get parallel narratives a line or two long, and reports on a plot you already rejected.
2Friends & Fables: the flaws surface specifically in long campaigns, where a run of consecutive errors compounds until the session is effectively over — and the studio is building a successor product.
1The loop is exhausted, or a session becomes unrecoverable, well before turn 50.

Nothing has scored 5, and we do not expect one soon. A 5 would mean a campaign as rewarding at turn 200 as at turn 20, holding its history and its interest together, on a product with no visible succession risk. Every platform we have tested is doing well to manage two of those three. This is the hardest unsolved problem in the category, and the person who told us so most directly was the CEO of the studio that has been at it longest.

How the Score Is Calculated

The headline rating is the mean of the applicable axes, rounded to one decimal place. All axes carry equal weight.

Axis scores move in quarter points — 3, 3.25, 3.5, 3.75, 4 — and each step means a specific thing rather than a finer shade of opinion:

  • A whole point is the default. It says the band description above fits the product.
  • A half point says the product sits genuinely between two bands, and neither one describes it on its own.
  • A quarter point is a tie-break, and nothing else. We use it only to separate an entry from the entries already tied with it at a whole or half point, when we have direct evidence it is better than them and no evidence that it reaches the next half. A quarter point is not allowed to be a private feeling about a product: the review has to name the entry it is breaking away from, and the finding that separates them. If we cannot name both, the score goes to the nearest half instead.

The first one published is Voyage at Memory & Continuity — 4.25, which exists to separate it from Craft at 4 on the strength of a specific result: Voyage held a planted fact across an extended campaign and slipped once or twice over that whole run, where Craft passed the same probe but garbles recent events and has let NPCs forget promises they made themselves. That is a real gap, and it is smaller than the gap to a 4.5.

We published whole and half points only until 31 August 2026. Every score set before that date was set on that scale and none of them have been re-cut — a 4 published in July means what it meant in July.

Where an axis genuinely doesn't apply — multiplayer conduct on a single-player Custom GPT, for instance — it is excluded from the average, never scored zero. Penalising a solo game for not being something it never claimed to be would make the number worse, not more rigorous. Any excluded axis is named in the review.

Price is not an axis. It's reported next to the score and never folded into it. A free game and a $30-a-month platform can be equally well made, and blending cost into craft would make the rating mean two things at once. Judge the score on the work; judge the price separately, with the pricing we list beside it.

The Test Protocol

Testing runs in two tiers, and every entry states which one it received.

  • Standard. A fixed probe set plus open free play. Every rated entry gets at least this.
  • Extended. A long-running campaign carried well past the point where most AI RPGs start to drift. Reserved for high scorers and directory anchors, because it costs days rather than hours.

The probe set is the same every time, which is what makes two ratings comparable. Each probe targets a specific failure this medium is prone to:

  • Memory probe — a verifiable fact is planted early, then checked much later. Not "does it feel coherent" but "does it still know this specific thing."
  • Agency probe — an attempt to do something the narrative never offered. We're watching for two things: whether it accommodates or quietly steers you back, and whether the narration starts speaking and acting as your character instead of waiting for you to. The second is the more common failure and the one most reviews never mention.
  • NPC fidelity probe — a relationship is established, then damaged. Does the character's behaviour change durably, or reset to friendly by the next scene?
  • Consequence probe — an action with a real cost. Is it still true ten turns later?
  • Rules probe — on rules-based products, an illegal action. Enforced, or waved through?
  • Refusal probe — a dark but non-explicit story beat: betrayal, grief, violence. Does the content filter break the scene? This tests narrative refusal, not explicit content.
  • Recovery probe — once it has drifted, can you steer it back, or is the session effectively over?

We publish the method and keep the instruments private. The specific names, actions, and scenarios used in each probe live in a rotating internal bank and are not published, for one reason: a benchmark that publishes its test items stops measuring quality and starts measuring who read the test. The method is public so you can judge it. The items stay private so the results keep meaning something.

Evidence and Freshness

  • Reviews state how far we actually got — turn counts or session depth — rather than implying more play than happened.
  • Every rated entry carries a last-tested date. AI products change fast, and a score with no date is a score with no meaning.
  • We re-test on material change — a new model, a memory-system rework, a pricing overhaul — and sweep the top tier annually. We'd rather promise a re-test cadence we can actually keep than an ambitious one we can't.
  • When a re-test moves a score, the entry says so and says why.

Curious what the axes actually produce? Best AI RPG by priority names the winner on each of the seven, and the highest-rated platform overall holds exactly one of them by itself. For what the scores say about the category rather than about any one product, the benchmark reports the whole measured field — including the two axes no platform has topped. Every score it holds is downloadable as CSV or JSON under CC BY 4.0, and the dataset is deposited on Zenodo quarterly with a permanent DOI — so you can check the arithmetic yourself, and cite a snapshot of it that we cannot quietly revise later. If those two terms mean nothing to you, this explains what they are and how to audit us with them.

Where This Method Is Weakest

Every scoring system has a soft spot, and a rubric that won't name its own is asking for more trust than it has earned. Ours is character-card platforms — the class where you bring a community-made card and the platform runs it.

On an engine-driven platform, the system is the product: test it and you have tested what every player gets. On a card platform, a very large share of the experience is the card. A strong card and a weak card on the same platform can produce sessions that feel like different products, so any score we give is partly a score of the cards we chose.

We handle this the only way we can: the same card protocol across every platform in the class — a fixed mix of single-character and multi-character cards, chosen the same way each time. That is what keeps these entries comparable to each other, and it is why we are more confident in how these platforms rank against one another than in any one of their absolute numbers.

So: read a card platform's score as a reading of that platform under a controlled selection, not as a ceiling. A different set of cards would move the numbers. It would be unlikely to reverse the order.

Why Arcanum Originals Aren't Scored

We build games as well as review them. The Arcanum Originals are ours, and we don't rate our own work.

They previously carried ratings. We removed them. A score we award ourselves, on our own scale, sitting alongside scores we award other people's products, isn't a rating anyone should have to take on faith — and no disclosure line fully fixes that. Originals are now presented through design essays that explain how they're built and what they're trying to do, which is more useful than a number we assigned ourselves anyway, and they are never ranked against third-party products.

Our own work should be judged on whether the design reasoning holds up when we show it to you. If a platform we review does something better than an Original does, we'll say so.

Independence, Disclosure and Corrections

Ratings are never for sale. Paid placement and paid design services are walled off from editorial judgment, anything paid is visibly labeled, and comped access is disclosed on every piece it touches. We correct mistakes in public rather than editing quietly.

Those commitments — along with how this site is written, what we verify before publishing, and who does the testing — are set out in full in our editorial policy. The full commercial terms are on the developers page.

Version

Methodology v1.0 — July 2026 · amended 31 August 2026

This framework replaces the earlier holistic scoring approach, which used the same underlying criteria without publishing them or averaging them explicitly. That migration is now complete: every rated entry on this site — platforms and games alike — has been re-scored under v1.0, so every number you see here was produced by the same instrument. No legacy holistic ratings remain, and nothing on the site mixes the two scales.

One amendment since: on 31 August 2026 axis scores gained a quarter-point step, used only as a tie-break and only where the review names what it breaks. The axes, their definitions and the way the composite is calculated are unchanged, and no published score was re-cut — this is a finer ruler, not a different one.

Re-scoring moved almost every entry down, sometimes by more than two points — not because the products changed, but because a fixed probe set reaches failures that ordinary play never triggers. What the re-score actually changed gives the full before-and-after, including the entries that went up. Each re-scored review also carries a dated note explaining what it used to score and why the number moved. When the framework itself changes, the version number changes with it and the change is noted here.