Top

What Changed in AI RPGs: August 2026

Six platforms earned their first rating in August, the top of our board changed hands twice in the last seven days of the month, and the thing most people were waiting for arrived with about twelve hours to spare.

Correction (2 September 2026). This note said in four places that Voyage was Android-only and that an iPhone player had nothing to install, including in the table below and in the list of things we were still waiting for. That was wrong when we published it: Voyage AI RPG Platform, published by Latitude Inc., has been on the App Store since 26 August 2026. We read the Google Play listing and Latitude’s announcements and never opened the App Store listing itself. The false lines are corrected here rather than left standing, because the rule against editing a dated record protects what was true on the day — not what was already wrong on it.

This is a monthly record of what actually moved in the AI RPG field, written from our own testing rather than from press releases. It is dated on purpose: everything below was true on 31 August 2026, and we do not go back and edit it later. Scores come from the benchmark, which now measures 21 platforms across seven axes.

August was the busiest month this site has had, and almost all of it landed in the final week.

Six Platforms Scored for the First Time

PlatformScoreHow it chargesThe short version
Voyage4.4/5One Latitude sub, $14.99–$99.99/moThe highest score we have awarded. Both phone stores
Craft4.3/5Free tier, $19.99–$99.99/moWorldbuilders write their own rules as code
Tidefall3.3/5Free tier, $14.99/moHand-authored world; the widest spread on our board
Tabled3.1/5$14.99 onceServer-audited dice, asynchronous play for six
WyrdTale3.0/5FreeThe game engine runs outside the model
ArcQuill2.6/5Credits, $9.99–$54.99/moPublishes a reliability score for each of its own models

Voyage is the month’s real event, and it arrived on the 31st. It entered at 4.4/5 — the highest composite we have awarded anything — and it did it by making consequences real: hit points, positioning and permanent death are game state rather than facts a model has to remember. Walk a party into a fight it has not prepared for and the party is wiped, with no negotiating afterwards. That single behaviour is most of why it scores where it does. Everything you achieve in it is something the engine could have refused you.

One thing to know before you plan around it: inventory and equipment still misbehave often enough over a long campaign that it cost the platform a point rather than a footnote. This paragraph also carried a second thing, and it was wrong — see the correction at the top.

Craft took the top of the board six days earlier, at 4.3/5, and held it until Voyage passed it by a tenth. It is the first thing we scored above 4, and it got there by refusing to be a game: it ships no house ruleset and hands the design surface to its players, who write new gameplay systems as code. Nothing else in the directory does that. What it charges is time — turns take long enough to process that we spent about as much of a session watching one resolve as playing it.

The gap between those two is worth stating plainly, because it is the smallest one on the board and it is the reason our scale changed this month (below). Craft is the more capable tool. Voyage is the better game.

Tidefall came in at 3.3/5 and is the strangest card we have filled in — three 5s and two 1s, the widest spread of any entry. Everything in it is made by hand first and performed by a model second: the quests are written by people, the illustrations drawn, the soundtrack scored. That buys characters who genuinely refuse you and a world that does not drift. It costs you the ability to leave the story, or to carry an inventory, or to see a currency.

Tabled entered at 3.1/5 by taking the dice out of the model’s hands entirely — every roll is made and audited on a server and handed to the game master already committed. Two chapters into our campaign we decided the party’s quest was not our character’s problem and simply left, and Tabled let us go rather than folding the plot around to meet us. It is not a clean sweep: it scores Determinism & Fairness — 2, because the plot we walked away from carried on without us and still paid our character experience for fights we were not present at. Honest dice, dishonest story.

WyrdTale scored 3.0/5 and charges nothing — no credits, no metered turns. Read the small print before calling it free: you connect your own Claude or ChatGPT subscription, so the cost moves rather than disappears.

ArcQuill came in at 2.6/5, and it does one thing we would like the rest of the category to copy: it publishes a reliability percentage and a credit cost for every model on its roster, including the ones it rates poorly.

What Didn’t Change: The Three Hard Axes

Six new platforms, and the shape of the board is almost the same. Across all 33 scored products, nobody has scored a 5 on memory, player agency, or longevity.

Almost the same, because one of the three finally moved. Memory & Continuity had capped at 4 since we started publishing axis scores, and Voyage broke it — not by much, but genuinely: 4.25, the only score above 4 anywhere on that axis. It held a fact we planted early across an extended campaign and slipped once or twice over the whole run. Player agency and longevity both still cap at 4, untouched.

The field averages tell the same story they told in July. Across every product we have scored, memory averages 2.66, player agency 2.36 and longevity 2.33 — the three lowest of the seven. Where 5s exist, they are on the more tractable problems: NPC fidelity, mechanical depth, determinism, signature design.

Every top score in this category is still on the easier problems. Six new entries did not dent that, and one of them is the highest-scoring thing we have ever rated.

The app stores told us the same thing from the other end. We audited every AI RPG app this month and found only eight of the twenty-one launched platforms ship one at all. The apps with enough ratings to mean anything sit above 4 stars with players while averaging well below that on our board. That is not the stores being wrong. A store rating is overwhelmingly a first-session verdict: did it open, did it look good, did the first scene land. Our axes are weighted toward what breaks at turn fifty.

The Scale Gained a Quarter Point

A methodology change, because the top of the board got crowded enough to need one.

Axis scores have used whole and half points since July. On 31 August we added quarter points, and they are legal in exactly one situation: as a named tie-break. Voyage’s Memory & Continuity — 4.25 is the first and so far only one. It exists to separate Voyage from Craft at 4, which passes the same memory probe but garbles recent events and lets NPCs forget promises they made themselves. That is a real gap, and it is smaller than the gap to a 4.5.

The rule that makes this honest rather than convenient: a review using a quarter point must name the entry it is breaking away from and the finding that separates them, or the score goes to the nearest half. Nothing finer than 0.25 exists. No published score was re-cut when the step was added. The full rule is at how the scale works.

We also explained, at some length and for the first time in plain English, what the DOI on our benchmark actually is — what a permanent archived snapshot does, what it does not do, and how to use it to check our arithmetic or catch us revising a score quietly. It is written for people who have never published a paper, which is nearly everyone who reads this site.

Listed, Not Yet Scored

Dungeons Deep joined the directory as an unrated entry after a 24-hour first look in its closed beta. It inverts the category’s usual bet: the adventures are written by humans and the AI is the referee rather than the author. The refereeing is genuinely tight — rules held, loot mattered, the shops worked as an economy. The cost is the cast, since one voice delivers the narration, the combat and every NPC in the same register.

We do not score pre-release software. That rule cost Voyage a score for four months, and stopped costing it one on 31 August.

What Developers Told Us

Two developers answered questions on the record in August.

Nick Walton, Latitude’s CEO, told us on 10 August that Voyage’s open beta was “only weeks away” and put full 1.0 at “early 2027.” The beta opened on schedule, three weeks later. The 1.0 date stands and is the thing to watch next.

Eric, founder of Mutiny Labs, put MythEngyn’s 1.0 at November, with the beta label tied to exactly one component: the rules engine. He also reported the longest campaign on the platform at 4,288 turns, a single continuous story. That figure is developer-reported and we have not verified it — but even discounted heavily it sits an order of magnitude past our own published reference point for where campaigns start coming apart. If a number like that survives our own testing, our longevity axis needs a harder ceiling rather than a higher score.

That is now the second time in two months a developer has told us our longevity axis is under-specified. Two independent people arriving at the same criticism is not a coincidence, and it is a promise we have made in public to go and fix.

Still Waiting

Voyage came off this list. Four things replace it.

Craft’s multiplayer. Cofounder Will, on 25 August: “we’re going to be rolling out multiplayer project editing within the next few days and then multiplayer gameplay sometime within the next month hopefully.” Neither had shipped when we tested, so nothing about group play on Craft is rated. It is the second-highest-scored platform in the directory and it is currently single-player only.

MythEngyn’s 1.0, targeted at November, gated on the rules engine.

Voyage on iOS. There is no App Store listing at all as of 31 August. Wrong when we published it: the listing went up on 26 August. See the correction at the top of this note.

Voyage’s 1.0, early 2027 on Latitude’s own account.

Corrections We Published

The unglamorous half of a monthly record. Everything here was wrong on this site until we fixed it in August:

  • Voyage is not sold separately. One Latitude subscription covers both Voyage and AI Dungeon. We had described Voyage-only pricing, including the unadvertised Shadow tiers, which also cover both.
  • AI Dungeon’s context allowance is per-model, not per-tier. The number on the pricing page is a ceiling that several models never reach and one exceeds. We replaced our tier-level claims with the full per-model table across all five tiers.
  • Credits can buy context, not just images. Latitude’s help centre states subscribers can spend one credit per action for extra context on some models. That makes memory a running cost on AI Dungeon, and we had never documented it.
  • Our model guide listed two models that no longer exist and was missing four that do.
  • Voyage’s Shadow tiers quietly halved a model’s context mid-month. The second model on every Shadow line read Raven at 32k, 64k and 128k when we ran the Walton interview on 10 August. It now reads GLM 5.2 at 16k, 32k and 64k — half the context on that slot at every tier, with prices and memory allowances unchanged. We corrected the directory entries and dated the figures in the interview rather than rewriting them.
  • Tidefall’s price was wrong the day we published it. We had its Unlimited tier at $17 a month. It is $14.99, billed monthly, with Paddle as merchant of record. Fixed within hours, and it is the reason we now check a price against the checkout rather than the pricing page.
  • Tabled’s iOS app was already live when we said it wasn’t. Our review and its directory entry both said iOS was “in App Store review.” It had listed on 15 August 2026, six days before we published that sentence.
  • Janitor AI ships its own apps. Our entry described mobile play through the browser as the only option. There has been a Janitor iOS app since November 2025, with an Android build alongside it.
  • Two rounds of stale superlatives, from our own scoring passes. Voyage taking Memory outright at 4.25 left eight published sentences still saying the axis caps at 4, and Voyage and Craft taking the top of the board left nine more calling some other platform “top-rated,” “joint-best” or “the highest-rated after.” Several were in structured data rather than prose, which is where this class of error hides. All seventeen are corrected.

A note on where these come from, because it matters more than the corrections themselves. The per-model context table exists because a reader called our AI Dungeon review “completely inaccurate” and declined to say why, so we went and measured it ourselves — and ended up publishing the best original data we produced all month. Latitude’s CEO sent notes on his own interview after publication, and three quotes were corrected. The rest we found by re-checking our own pages against vendor documentation, which is a thing worth doing on a schedule rather than when challenged.

The last item on that list is the one worth dwelling on. A score changing is not one edit; it is every sentence anywhere on the site that described the old order. We have now been caught by that four scoring passes in a row, each time by the previous pass’s leftovers. A review site with no corrections is not more accurate than one with a list this long. It is just quieter about it.

If You’re Picking Something This Month

If you want the best thing we have played, that is Voyage at 4.4/5 — on either phone store, with a Latitude subscription that gets you AI Dungeon alongside it. If you would rather not install anything, Craft at 4.3/5 is the closest thing — it installs to a home screen from the browser and has a real free tier, though it is single-player until multiplayer lands.

If you play with friends and cannot get everyone in a room, Tabled is still the strongest pick and costs $14.99 once for a table of six. If you want an authored, cinematic campaign rather than an open sandbox, Tidefall is unlike anything else here — read its card first, because what it takes away is as distinctive as what it gives. If you want to spend nothing at all, Janitor AI scores 3.0/5 with a genuinely free tier, and WyrdTale matches that score while expecting you to bring your own Claude or ChatGPT subscription.

If you are weighing a subscription anywhere, what an AI RPG actually costs covers the five ways platforms meter play — and why the shape of the meter matters more than the price. The highest-scored platform on the board can be played for nothing.

This note is published monthly. It records what changed on our board, what developers told us, and what we got wrong — dated, and left standing as written.