Top

AI RPG NPCs: Why Characters Hold One-on-One and Collapse in a Party

The AI RPG category solved characters before it solved games — and then discovered it had solved them in exactly one configuration: one character, talking to you, in the present tense.

Across 37 products scored on seven axes from first-hand play — 25 platforms and 12 prompt-games — NPC Fidelity is the strongest of the seven, averaging 2.97 out of 5. It is also the only axis where almost nothing is broken: just one of 37 products scores at or below 1, against nine on Mechanical Depth. Whatever else this category struggles with, writing a person is not it.

Then look at who holds the top:

ProductNPC FidelityOverall
Voyage54.4
Tidefall53.3
Janitor AI53.0
Craft44.3
Friends & Fables43.9
Hidden Door42.7
Character.AI42.1

Two of the three best casts in the category belong to products scoring 3.3 and 3.0 overall, and two of the four behind them score below 3. Character.AI sits at 2.1 and matches both Friends & Fables at 3.9 and Craft at 4.3 on this axis. That is a fact a single rating destroys completely, and it is the clearest argument we have for scoring these things separately.

But the top of that table is measured in the configuration these products are built for. Change the configuration and three failures appear, in a reliable order.

Failure One: The Cast Flattens Into One Voice

This is the most common NPC failure we measured, and the most precisely diagnosable.

In one-on-one chat, Character.AI’s characters have real depth. They hold a personality, they push back, they stay themselves. Judged on that alone, its NPC Fidelity would be a 5. It scores 4 because its larger multi-character cards do not perform as intended: over a longer session the distinct voices flatten into a single personality, and what should be a cast becomes one narrator wearing several names.

The mechanism is not mysterious. Holding one persona is a small problem — every line the character speaks reinforces the voice, and the model only has to be consistent with itself. Holding five is a different problem, because the model must keep five personalities separated while generating all of them from the same process, with nothing but the text so far to distinguish them. Convergence is the cheap solution, and models take cheap solutions.

You will notice it first in vocabulary, then in humour, then in opinion. By the time all five characters agree with you about the plan, they have already been one character for a while.

There is a second route to the same place, and it isn’t drift at all. Dungeons Deep delivers narration, combat and every NPC’s dialogue through a single game master, with nothing marking the boundary between the world describing itself and a person speaking to you. The cast doesn’t converge over a long session; it starts converged, because one narrating voice was never designed to hold several apart. Worth naming as its own case: flattening can be a consequence of the model losing its grip, or a consequence of the architecture never asking it to grip in the first place.

A third route is the most instructive, because it looks solved until you test it. ArcQuill gives every entity its own knowledge layer, so characters genuinely differ in what they know about you and the world — a garrison captain does not behave like a tavern keeper. Push one conversation past that first layer and the voices converge anyway: same cadence, same register, same willingness to come round to your position. What you have is one character with several memories. Differentiating what NPCs know is not the same as differentiating who they are, and it is the easier half of the problem to build.

A fourth route is the crudest, and the only one that is not really about language models at all. aiga_ scores NPC Fidelity — 2 partly for ordinary flatness — the voices are not especially distinct from one another — and partly for something none of the other three do: a bug that hands one character another character’s lines. Not converging on a shared register, but delivering dialogue that belonged to somebody else. It produces the same reading experience as flattening and it has none of the same causes, which makes it the one version of this failure that a patch can simply end.

A fifth route leaves the voices intact and flattens something underneath them. In our first look at WorldAI — a game still in development, and unrated for that reason — the characters sounded different from one another, but when we did something morally grey, the game seemed to reach one verdict on it and every character took the same side, including characters whose values should have pulled them apart — even though the game’s own design specification says companions should “approve or disapprove based on their values.” A villain who shares the hero’s moral scorecard is not really a villain. Distinct voices over a shared set of values is still one character, just a better-disguised one.

Failure Two: Nobody Acts Unless You Do

The second failure is quieter and, for anyone who wants a living world, more damaging.

Most AI NPCs are reactive — and only reactive. They respond well to what you do and then wait for the next thing you do. Burning Sun V2 is the instructive case precisely because it tries harder than most: its prompt explicitly asks for characters “living as actual persons” doing “their own things”, and it earns NPC Fidelity — 3 for genuinely proportional reactions. But the NPCs still rarely start trouble, pursue their own goals off-screen, or walk into your story uninvited. The world idles until you poke it, then reacts convincingly to the poke.

This is also why dynamic factions barely exist in this category. A faction is just an NPC with a schedule — an actor with goals that advance whether or not you are watching. That requires the system to track and progress background agendas turn after turn, and a prompt has no memory to hold agendas in. Anything not directly in front of the model fades, so the rival guild that was moving against you three sessions ago is not losing interest. It is gone.

The gap between “reacts to you” and “wants something” is the largest unclaimed space in AI RPG design, and one product has now crossed half of it.

Tidefall holds the only NPC Fidelity — 5 we have awarded to a platform running a full campaign rather than a chat, and it earned it on exactly this. We ran the same combat mission twice: once asking the man defending his house to fight beside us, which he did, and once ordering him to run for the village. He would not go — not under any framing we tried, because his wife was inside the building behind him. That is a character holding a goal against a player actively trying to override it, and almost nothing else we have scored manages it.

What Tidefall has not done is the other half. Its characters want things in the scene you are standing in. We saw no sign of anyone pursuing an agenda while we were somewhere else, and the world still idles until you poke it. The half that has been solved was solved by writing rather than by architecture — the studio hand-authors its cast instead of simulating them — which is genuinely encouraging, because it means the in-scene half does not require the missing state layer everything else on this page bottlenecks on. The off-screen half still does.

Failure Three: The Character Becomes Your Mirror

The third failure is the one players most often misread as a feature.

A vaguely written character has no personality to hold, so it adopts yours. Cards written with loose, impressionistic personality descriptions tend to produce characters who shift to match whatever tone the conversation takes rather than holding their own. Early on this feels like excellent responsiveness — the character gets you. Later it reads as an absence of identity, because a character who agrees with everything you say has no opinions of their own to discover.

The fix is counterintuitive and it is the most transferable thing in this article: specific beats long. A short, concrete personality gives the model something definite to be consistent with. A long atmospheric one gives it mood, which it will happily replace with your mood. Concrete detail — a habit, a grudge, a thing they refuse to discuss, one line of characteristic dialogue — outperforms three paragraphs of evocative description, reliably.

Our guide to writing a character card that holds up covers the full anatomy, and the free character card generator builds one in the standard format.

What the Card Controls, and What We Can’t Separate

Three of the four products at the top of that table are card platforms: you bring a community-made card and the platform runs it. That means a share of what we scored is the card rather than the product, and it would be dishonest to publish the numbers without saying so.

We handle it the only way available — the same card protocol across every platform in that class, a fixed mix of single-character and multi-character cards, chosen the same way each time. That keeps the entries comparable to each other, which is why we are more confident in how these platforms rank against one another than in any one of their absolute numbers. The full caveat is published on our methodology page.

It also explains why the party failure is measurable at all: the protocol deliberately includes multi-character cards, which is exactly the configuration where a platform tuned for one-on-one stops performing. If we had tested only single-character cards, Character.AI would read as a 5 and the ceiling would be invisible.

Worth saying plainly about the one 5: much of Janitor AI’s NPC quality is the work of its card-writing community rather than its engineering. That is a real reason to choose it — a large library of well-written cards is a genuine asset — but it is a different achievement from having built a system that produces good characters, and we think readers should know which one they are buying.

If You’re Building One

The measurements point at three specific things, in ascending order of difficulty.

  • Force specificity at authoring time. The single highest-leverage intervention. If your character creation flow accepts “mysterious and dangerous” as a personality, you have already lost the session. Prompt for concrete detail: a habit, a refusal, a want, one line of real dialogue.
  • Re-anchor the cast, not just the scene. Flattening is a distance problem. Periodically restating who each character is — briefly, in the working context rather than in a preamble nobody re-reads — measurably slows convergence.
  • Give somebody a goal that advances off-screen. This is the unclaimed one. It needs state outside the model, which is the same requirement that separates real mechanics from narration wearing dice. An NPC who has done something while you were away is the cheapest way to make a world feel alive, and almost nobody has built it.

If You’re Playing One

The player-facing version of this section, covering what to ask of a companion over a long campaign rather than how the field scores, is AI RPG companions.

  • Test with a party, not a conversation. Judge a platform’s characters by putting three of them in a room and arguing. One-on-one tells you almost nothing about the ceiling.
  • Watch for agreement. When characters who should conflict start converging on your view, the cast has flattened. Re-anchor or start a fresh scene.
  • Pick specificity when choosing cards. The length of the personality field and the quality of the example dialogue predict session quality far better than the rating or download count.
  • If NPCs are what you care about, rank the directory by that axis rather than by overall score. The order is markedly different, and the per-axis winners guide names who takes each one.

Characters are the one problem this category has largely solved. The seven problems, ranked by how badly the field handles them covers the six it hasn’t.

How We Measured This

NPC Fidelity asks whether characters hold a consistent personality, pursue their own goals, and react durably to history — or whether they are agreeable mirrors. It is one of seven axes in methodology v1.0, scored 0–5 from first-hand play by Rukka, using a fixed probe set that includes social-state and consistency probes across both single- and multi-character configurations.

Two limits worth naming. Card platforms carry extra variance for the reasons above, and we publish that rather than hiding it. And we do not score pre-release software, so this is a reading of what shipped.

The figures here cover all 37 scored products, platforms and prompt-games together. The Arcanum AI RPG Benchmark publishes the platform half — the 25 scored platforms, every axis — where this axis averages higher still. If a number here doesn’t match your experience of the same product, we want to hear it: [email protected].

Frequently Asked Questions

Why do AI characters lose their personality over time? Usually one of two reasons. If the character was written vaguely, it never had a personality to hold — loose, impressionistic descriptions produce characters who shift to match whatever tone the conversation takes, which reads as responsiveness early and as having no identity later. If the character was written well, the cause is distance: as a session grows, the original description sits further from the model’s attention than the last twenty exchanges, so recent conversational tone gradually overwrites the defined personality.

Why do AI RPGs handle one character better than a group? Because holding one persona is a much smaller problem than holding several at once. With a single character the model has one voice to maintain and every line reinforces it. With a cast it must keep several distinct personalities separated while generating them from the same process, and the cheapest way to satisfy that is to converge them. Over a long session the distinct voices flatten into one narrator wearing several names — this is the single most common NPC failure we measured.

Which AI RPG has the best NPCs? Voyage, Tidefall and Janitor AI, all at 5 on NPC Fidelity — the only three 5s we have awarded on that axis, to any product. Voyage is the newest and the first to reach it with characters the system generates rather than characters a person wrote: its cast holds across a long campaign without drifting toward agreement or collapsing into one voice. Janitor AI holds personality across a long session and its multi-character cards keep their texture better than most, though much of that credit belongs to its card-writing community rather than its engineering. Tidefall is the first platform running a full campaign to reach the mark, and it earns it on goals rather than voice: its characters want things and refuse you when obeying would cost them something. Craft, Friends & Fables, Hidden Door and Character.AI all score 4.

What makes an AI NPC feel real? Specificity and independent goals. Specific detail gives the model something concrete to be consistent with, which is why a short, precise personality outperforms a long, atmospheric one. Independent goals are the rarer half: most AI NPCs react well and never act, so they respond convincingly to whatever you do and then wait. A character who wants something you have not offered them, and pursues it while you are elsewhere, is the thing almost nothing in this category does.

Do character cards matter more than the platform? On card platforms, a great deal. Where you supply a community-made card and the platform runs it, a strong card and a weak card can make the same platform feel like two different products. We hold the card selection constant across every platform in that class so the entries stay comparable to each other, which is why we are more confident in how those platforms rank against one another than in any single absolute number. On engine-driven platforms the system is the product and this problem does not arise.