Top

// benchmark · prompt-games

The Prompt-Game Benchmark

2026 H2 edition · 12 games measured · last updated 2026-09-18 · the platform board is here

Six of the seven axes cap at 3 across every prompt-game we have scored. Exactly one escapes: Signature Design, which reaches 4. A prompt can be original. It cannot be deep.

These are prompt-games: a text file, a Custom GPT, or a project you load into a model you already pay for. No server, no engine, no database — the model is the whole product. We score them on the same seven axes as the hosted platforms, because a player's evening does not care which architecture produced it.

This is the sibling of the platform benchmark, kept separate on purpose. Averaging a paste-in prompt against a hosted product with a rules engine produces a number that describes neither — and the platform dataset has said since it was first deposited that prompt-games are measured elsewhere. This is that elsewhere.

What Is In It — and What Isn't

Everything here is a third-party game we played first-hand and scored. Anything still unpublished or unreachable is not on the board, on the same rule the platform side uses: we score what you can actually load and run.

Deliberately absent (4): Aevum Realm Architect, Eirathis Strider, Star Freighter Drift and The Chronicler. These are the Arcanum Originals — our own games, and we never rate our own work. That is a policy, not a gap: they are excluded because scoring them would make this board worthless, not because they are unfinished. Several are free, and you can judge them yourself.

Scores run 0–5 per axis; the headline number is the mean of the axes that apply. The rubric, the scale and the test protocol are all on the methodology page.

Finding 1 — The Ceiling Is Structural

Across 12 games and seven axes, almost nothing reaches past 3. Six axes stop dead at 3: Memory & Continuity, Player Agency, NPC Fidelity, Mechanical Depth, Determinism & Fairness and Longevity. The single exception is Signature Design, at 4 — held by Isekai RPG and 8-Bit Kingdoms.

That is not twelve authors independently running out of talent. It is the shape of the format. Memory, fairness and mechanical depth all require something to hold state and adjudicate outside the model — a database, a dice server, a rules engine. A prompt has none of those and cannot acquire them. It can ask the model to track a character sheet, and the model will, until it doesn't.

The one axis that escapes is the one that needs no engine at all. Signature Design measures whether the thing has a point of view — a premise, a voice, a reason to exist rather than the twentieth generic fantasy sandbox. That is pure authorship, and authorship is exactly what a prompt is made of.

Finding 2 — The Format Has No Floor

Signature Design is the only axis that clears 3, and it runs from 0 to 4 — the widest spread on the board. Three of the 12 games score 1 or below on it.

Best and worst living on the same axis is what a format with no floor looks like. A hosted platform ships a baseline: even a mediocre one has an interface, a memory system and somebody's product decisions holding it up. A prompt has whatever its author wrote and nothing else. Where the author had a real idea it shows immediately, and where they didn't there is no engine underneath to carry it — the result is a generic sandbox a two-line prompt would have matched.

The averages say the same thing more quietly. NPC Fidelity is the strongest axis at 2.42, and Mechanical Depth the weakest at 1.75, with 4 of 12 games at 1 or below. Both are within about a point of each other, which is the whole story of this board: not much separates anything.

Finding 3 — The Whole Board Fits Between 1.6 and 2.6

That is a range of 1.0. On the platform board the same seven axes produce a range of 3.0 — 3.0× wider. Compressed range is what a hard ceiling looks like from the outside: when the top of every axis is capped, the composite cannot separate anything very much, and the gap between the best prompt-game here and the worst is smaller than the gap between two mid-table platforms.

The practical reading: on this board, the ranking matters less than on the other one. Pick by premise. Anything in the top half will play about the same, and the thing that will actually decide your evening is whether you want a kingdom sim, a dungeon crawl or a character study.

Where the Field Stands, Axis by Axis

Mean is the average across all 12 measured games. Best is the highest any single game has reached.

AxisThe questionMeanBestHeld by
NPC Fidelity Consistent personality, own goals, durable reactions — or agreeable mirrors? 2.42 3 Solo RPG Master, Isekai RPG, Vantiel, Burning Sun V2, RPG GPT, Solo D&D Starter Prompt
Signature Design Does it do something nobody else does? 2.33 4 Isekai RPG, 8-Bit Kingdoms
Player Agency Railroading AND puppeting — can you act off-path, and does the narrator stay out of your character's mouth? 2.25 3 Solo RPG Master, Deep Saga, Vantiel
Memory & Continuity Does the world remember what happened, and for how long? 2.25 3 Solo RPG Master, Valkyrie's Biggest Gig, Gamekeeper RPG Prompt
Determinism & Fairness Outcomes earned and consistent, or arbitrary and luck-washed? 2.17 3 Deep Saga, 8-Bit Kingdoms, Classic Text Adventure Prompt
Longevity Does it survive past turn 50 and the novelty window? 2.00 3 Solo RPG Master, Vantiel
Mechanical Depth Real systems underneath, or narration wearing dice? 1.75 3 Isekai RPG, Valkyrie's Biggest Gig

The Board

All 12 measured games, every axis, sorted by composite. Read across the row rather than down the column — with a range this narrow, a single axis tells you far more than the average does.

GameScore MemoryPlayerNPCMechanicalDeterminismLongevitySignature
Solo RPG Master 2.6 3 3 3 2 2 3 2
Deep Saga 2.4 2 3 2 2 3 2 3
Isekai RPG 2.4 2 2 3 3 2 1 4
Vantiel 2.4 2 3 3 1 2 3 3
8-Bit Kingdoms 2.3 2 2 2 2 3 1 4
Valkyrie's Biggest Gig 2.3 3 2 2 3 1 2 3
Burning Sun V2 2.1 2 2 3 1 2 2 3
Gamekeeper RPG Prompt 2.1 3 2 2 2 2 2 2
RPG GPT 2.0 2 2 3 2 2 2 1
Classic Text Adventure Prompt 1.9 2 2 1 1 3 2 2
Realm & Companion RPG Prompt 1.9 2 2 2 2 2 2 1
Solo D&D Starter Prompt 1.6 2 2 3 0 2 2 0

Highlighted cells are the best score on that axis. Column headings are shortened — full axis names are in the table above.

Take the Data

The whole board, in full: download the CSV or the JSON. Both are generated from the same source as the table above at build time, so the file and the page cannot drift apart.

Published under CC BY 4.0, the same licence as the platform board. Reuse it, chart it, argue with it. The one condition is attribution with a link back.

This board does not have a DOI, and we are not going to pretend otherwise. The platform benchmark is deposited in Zenodo with a permanent archived snapshot; this one is not, yet. That means the numbers here can move without leaving a citable record of what they used to be — so if you are quoting this board in something that needs to stay checkable, quote it with the date you read it. What a DOI actually is explains the difference and why it matters.

How to Read This Honestly

Four limits, because a board this small should be upfront about what it can support.

  • 12 games is a small field. The platform board carries enough entries for its averages to mean something. This one is close to the point where a single new entry moves a field average visibly — treat the per-axis means as indicative, not settled.
  • A prompt-game's score is not a verdict on the model running it. These games run on different assistants, and we have not held the model constant across them. The number measures the game as we played it, not the assistant underneath.
  • Low here does not mean bad value. Every game on this board is free or near-free and runs inside a subscription you already have. A 2.4 that costs nothing and a 2.4 that costs $30 a month are not the same product, and the composite cannot see the difference. What an AI RPG actually costs covers that properly.
  • Prompts change without telling you. An author can rewrite a Custom GPT overnight and the link stays the same. Platform versions at least announce themselves; a prompt-game is whatever it was the day we played it.

Where to Go From Here

To pick something rather than read about the field: the games directory has all of them with what each one is for, and the Arcanum Originals are our own, free, and unrated by policy. If you would rather have an engine underneath, the platform board measures the products that ship one — and it is a different field, with a ceiling more than a point higher.

Corrections

If a number here doesn't survive contact with your experience of the same game, we want to hear it — [email protected]. Disagreeing with a verdict isn't a correction, but a score whose reasoning doesn't hold up is.