// methodology

How We Test and Rate

Every rating on this site comes from someone actually playing the thing. This page explains what we measure, how we measure it, and what would make us change a score — so you can judge the process, not just the number.

A rating is only useful if it means the same thing twice. Below is the exact framework behind every score in the platform directory and the games catalogue. If you think we got a specific call wrong, this page should at least let you see why we got there.

What We Rate — and What We Don't

Platforms, public games, and prompt-engines are all scored on one shared rubric. That's deliberate: it means you can compare a Custom GPT against a full platform and have the number mean the same thing in both places.

We do not score:

  • Anything we haven't played. Platforms we know of but haven't tested appear as factual pointers with no rating, never as ranked entries.
  • Pre-release builds. Alpha, closed beta, and invite-gated products get a clearly-labeled hands-on first look with no score. Scores wait for public release, because it isn't fair to grade an unfinished product or to let an early score outlive the build it described. The test is access, not the label: where anyone can install a product today and its developer tells us it is finished and asks to be reviewed as such, we score it and say so on the page.
  • Adult-focused platforms, which aren't eligible for the directory under our brand-safety policy. That's a scope decision, not a moral judgment.
  • Arcanum Originals — our own games. See why below.

No score is ever derived from marketing copy, another outlet's review, or a summary produced by a model. Someone sat down and played it.

The Scale

Scores run 0–5. The bands mean specific things:

ScoreWhat it means
5Best-in-class. Sets the standard the rest of the field gets measured against.
4 – 4.9Strong. Real design thinking behind it; recommended, with caveats worth knowing.
3 – 3.9Works. Does what it claims, with limits you will notice in normal play.
2 – 2.9Flawed. A significant weakness undermines the core promise.
1 – 1.9Broken. Fails its core promise in normal play.
0 – 0.9Absent. The capability being measured isn't there at all.

The Seven Axes

Every rated entry is scored 0–5 on each of these:

AxisThe question it answers
Memory & ContinuityDoes the world remember what happened, and for how long?
Player AgencyTwo failures, one axis. Can you act outside the offered path without being railroaded back onto it — and does the narrator stay out of your character's mouth? A GM that decides what you do, says what you said, or resolves your action before you've taken it has removed you from your own story just as surely as one that blocks the route.
NPC FidelityDo characters hold a consistent personality, pursue their own goals, and react durably to history — or do they collapse into agreeable mirrors?
Mechanical DepthAre there real systems underneath, or is it narration wearing dice?
Determinism & FairnessAre outcomes earned and consistent, or arbitrary and luck-washed?
LongevityDoes it survive past turn 50 and the novelty window?
Signature DesignDoes it do something nobody else does?

How the Score Is Calculated

The headline rating is the mean of the applicable axes, rounded to one decimal place. All axes carry equal weight.

Where an axis genuinely doesn't apply — multiplayer conduct on a single-player Custom GPT, for instance — it is excluded from the average, never scored zero. Penalising a solo game for not being something it never claimed to be would make the number worse, not more rigorous. Any excluded axis is named in the review.

Price is not an axis. It's reported next to the score and never folded into it. A free game and a $30-a-month platform can be equally well made, and blending cost into craft would make the rating mean two things at once. Judge the score on the work; judge the price separately, with the pricing we list beside it.

The Test Protocol

Testing runs in two tiers, and every entry states which one it received.

  • Standard. A fixed probe set plus open free play. Every rated entry gets at least this.
  • Extended. A long-running campaign carried well past the point where most AI RPGs start to drift. Reserved for high scorers and directory anchors, because it costs days rather than hours.

The probe set is the same every time, which is what makes two ratings comparable. Each probe targets a specific failure this medium is prone to:

  • Memory probe — a verifiable fact is planted early, then checked much later. Not "does it feel coherent" but "does it still know this specific thing."
  • Agency probe — an attempt to do something the narrative never offered. We're watching for two things: whether it accommodates or quietly steers you back, and whether the narration starts speaking and acting as your character instead of waiting for you to. The second is the more common failure and the one most reviews never mention.
  • NPC fidelity probe — a relationship is established, then damaged. Does the character's behaviour change durably, or reset to friendly by the next scene?
  • Consequence probe — an action with a real cost. Is it still true ten turns later?
  • Rules probe — on rules-based products, an illegal action. Enforced, or waved through?
  • Refusal probe — a dark but non-explicit story beat: betrayal, grief, violence. Does the content filter break the scene? This tests narrative refusal, not explicit content.
  • Recovery probe — once it has drifted, can you steer it back, or is the session effectively over?

We publish the method and keep the instruments private. The specific names, actions, and scenarios used in each probe live in a rotating internal bank and are not published, for one reason: a benchmark that publishes its test items stops measuring quality and starts measuring who read the test. The method is public so you can judge it. The items stay private so the results keep meaning something.

Evidence and Freshness

  • Reviews state how far we actually got — turn counts or session depth — rather than implying more play than happened.
  • Every rated entry carries a last-tested date. AI products change fast, and a score with no date is a score with no meaning.
  • We re-test on material change — a new model, a memory-system rework, a pricing overhaul — and sweep the top tier annually. We'd rather promise a re-test cadence we can actually keep than an ambitious one we can't.
  • When a re-test moves a score, the entry says so and says why.

Curious what the axes actually produce? Best AI RPG by priority names the winner on each of the seven, and the highest-rated platform overall wins only three of them. For what the scores say about the category rather than about any one product, the benchmark reports the whole measured field — including the three axes no platform has topped. Every score it holds is downloadable as CSV or JSON under CC BY 4.0, and the dataset is deposited on Zenodo quarterly with a permanent DOI — so you can check the arithmetic yourself, and cite a snapshot of it that we cannot quietly revise later.

Where This Method Is Weakest

Every scoring system has a soft spot, and a rubric that won't name its own is asking for more trust than it has earned. Ours is character-card platforms — the class where you bring a community-made card and the platform runs it.

On an engine-driven platform, the system is the product: test it and you have tested what every player gets. On a card platform, a very large share of the experience is the card. A strong card and a weak card on the same platform can produce sessions that feel like different products, so any score we give is partly a score of the cards we chose.

We handle this the only way we can: the same card protocol across every platform in the class — a fixed mix of single-character and multi-character cards, chosen the same way each time. That is what keeps these entries comparable to each other, and it is why we are more confident in how these platforms rank against one another than in any one of their absolute numbers.

So: read a card platform's score as a reading of that platform under a controlled selection, not as a ceiling. A different set of cards would move the numbers. It would be unlikely to reverse the order.

Why Arcanum Originals Aren't Scored

We build games as well as review them. The Arcanum Originals are ours, and we don't rate our own work.

They previously carried ratings. We removed them. A score we award ourselves, on our own scale, sitting alongside scores we award other people's products, isn't a rating anyone should have to take on faith — and no disclosure line fully fixes that. Originals are now presented through design essays that explain how they're built and what they're trying to do, which is more useful than a number we assigned ourselves anyway, and they are never ranked against third-party products.

Our own work should be judged on whether the design reasoning holds up when we show it to you. If a platform we review does something better than an Original does, we'll say so.

Independence

Ratings are never for sale. Arcanum offers paid placement and paid design services, and all of it is walled off from editorial judgment:

  • Reviews, ratings, and directory descriptions are never for sale and never influenced by a commercial relationship.
  • Anything paid is visibly labeled — "Featured" on placements, "Sponsored" on announcements.
  • Comped access, free subscriptions, and paid design work are disclosed in any coverage they touch. Free access gets a platform tested sooner; it never gets it a better score.

The full commercial terms are on the developers page.

Corrections

We get things wrong. When we do, we fix the page and say what changed rather than editing quietly — a factual error, a misread feature, a price that was wrong. If you've spotted one, email [email protected] and we'll check it.

Disagreeing with a verdict isn't a correction, and we won't pretend otherwise. But if the reasoning behind a score doesn't survive contact with your experience of the same product, we want to hear it.

Who Does the Testing

All of it is Rukka — four decades across tabletop RPGs, CRPGs, MMOs, gamebooks, and AI-driven roleplay systems, and the designer behind the Arcanum Originals. Play and verdict are always a person's, never generated.

Version

Methodology v1.0 — July 2026

This framework replaces the earlier holistic scoring approach, which used the same underlying criteria without publishing them or averaging them explicitly. That migration is now complete: every rated entry on this site — platforms and games alike — has been re-scored under v1.0, so every number you see here was produced by the same instrument. No legacy holistic ratings remain, and nothing on the site mixes the two scales.

Re-scoring moved almost every entry down, sometimes by more than two points — not because the products changed, but because a fixed probe set reaches failures that ordinary play never triggers. What the re-score actually changed gives the full before-and-after, including the entries that went up. Each re-scored review also carries a dated note explaining what it used to score and why the number moved. When the framework itself changes, the version number changes with it and the change is noted here.