AI RPG NPCs: Why Characters Hold One-on-One and Collapse in a Party

The AI RPG category solved characters before it solved games — and then discovered it had solved them in exactly one configuration: one character, talking to you, in the present tense.

Across 28 products scored on seven axes from first-hand play — 16 platforms and 12 prompt-games — NPC Fidelity is the strongest of the seven, averaging 2.86 out of 5. It is also the only axis where almost nothing is broken: just one of 28 products scores at or below 1, against eight on Mechanical Depth. Whatever else this category struggles with, writing a person is not it.

Then look at who holds the top:

ProductNPC FidelityOverall
Janitor AI53.0
Friends & Fables43.9
Hidden Door42.7
Character.AI42.1

Three of the four best casts in the category belong to products that score below 3 overall. Character.AI sits at 2.1 and matches Friends & Fables at 3.9 on this axis. That is a fact a single rating destroys completely, and it is the clearest argument we have for scoring these things separately.

But the top of that table is measured in the configuration these products are built for. Change the configuration and three failures appear, in a reliable order.

Failure One: The Cast Flattens Into One Voice

This is the most common NPC failure we measured, and the most precisely diagnosable.

In one-on-one chat, Character.AI’s characters have real depth. They hold a personality, they push back, they stay themselves. Judged on that alone, its NPC Fidelity would be a 5. It scores 4 because its larger multi-character cards do not perform as intended: over a longer session the distinct voices flatten into a single personality, and what should be a cast becomes one narrator wearing several names.

The mechanism is not mysterious. Holding one persona is a small problem — every line the character speaks reinforces the voice, and the model only has to be consistent with itself. Holding five is a different problem, because the model must keep five personalities separated while generating all of them from the same process, with nothing but the text so far to distinguish them. Convergence is the cheap solution, and models take cheap solutions.

You will notice it first in vocabulary, then in humour, then in opinion. By the time all five characters agree with you about the plan, they have already been one character for a while.

Failure Two: Nobody Acts Unless You Do

The second failure is quieter and, for anyone who wants a living world, more damaging.

Most AI NPCs are reactive — and only reactive. They respond well to what you do and then wait for the next thing you do. Burning Sun V2 is the instructive case precisely because it tries harder than most: its prompt explicitly asks for characters “living as actual persons” doing “their own things”, and it earns NPC Fidelity — 3 for genuinely proportional reactions. But the NPCs still rarely start trouble, pursue their own goals off-screen, or walk into your story uninvited. The world idles until you poke it, then reacts convincingly to the poke.

This is also why dynamic factions barely exist in this category. A faction is just an NPC with a schedule — an actor with goals that advance whether or not you are watching. That requires the system to track and progress background agendas turn after turn, and a prompt has no memory to hold agendas in. Anything not directly in front of the model fades, so the rival guild that was moving against you three sessions ago is not losing interest. It is gone.

The gap between “reacts to you” and “wants something” is the largest unclaimed space in AI RPG design. Nothing we have scored convincingly crosses it.

Failure Three: The Character Becomes Your Mirror

The third failure is the one players most often misread as a feature.

A vaguely written character has no personality to hold, so it adopts yours. Cards written with loose, impressionistic personality descriptions tend to produce characters who shift to match whatever tone the conversation takes rather than holding their own. Early on this feels like excellent responsiveness — the character gets you. Later it reads as an absence of identity, because a character who agrees with everything you say has no opinions of their own to discover.

The fix is counterintuitive and it is the most transferable thing in this article: specific beats long. A short, concrete personality gives the model something definite to be consistent with. A long atmospheric one gives it mood, which it will happily replace with your mood. Concrete detail — a habit, a grudge, a thing they refuse to discuss, one line of characteristic dialogue — outperforms three paragraphs of evocative description, reliably.

Our guide to writing a character card that holds up covers the full anatomy, and the free character card generator builds one in the standard format.

What the Card Controls, and What We Can’t Separate

Three of the four products at the top of that table are card platforms: you bring a community-made card and the platform runs it. That means a share of what we scored is the card rather than the product, and it would be dishonest to publish the numbers without saying so.

We handle it the only way available — the same card protocol across every platform in that class, a fixed mix of single-character and multi-character cards, chosen the same way each time. That keeps the entries comparable to each other, which is why we are more confident in how these platforms rank against one another than in any one of their absolute numbers. The full caveat is published on our methodology page.

It also explains why the party failure is measurable at all: the protocol deliberately includes multi-character cards, which is exactly the configuration where a platform tuned for one-on-one stops performing. If we had tested only single-character cards, Character.AI would read as a 5 and the ceiling would be invisible.

Worth saying plainly about the one 5: much of Janitor AI’s NPC quality is the work of its card-writing community rather than its engineering. That is a real reason to choose it — a large library of well-written cards is a genuine asset — but it is a different achievement from having built a system that produces good characters, and we think readers should know which one they are buying.

If You’re Building One

The measurements point at three specific things, in ascending order of difficulty.

  • Force specificity at authoring time. The single highest-leverage intervention. If your character creation flow accepts “mysterious and dangerous” as a personality, you have already lost the session. Prompt for concrete detail: a habit, a refusal, a want, one line of real dialogue.
  • Re-anchor the cast, not just the scene. Flattening is a distance problem. Periodically restating who each character is — briefly, in the working context rather than in a preamble nobody re-reads — measurably slows convergence.
  • Give somebody a goal that advances off-screen. This is the unclaimed one. It needs state outside the model, which is the same requirement that separates real mechanics from narration wearing dice. An NPC who has done something while you were away is the cheapest way to make a world feel alive, and almost nobody has built it.

If You’re Playing One

  • Test with a party, not a conversation. Judge a platform’s characters by putting three of them in a room and arguing. One-on-one tells you almost nothing about the ceiling.
  • Watch for agreement. When characters who should conflict start converging on your view, the cast has flattened. Re-anchor or start a fresh scene.
  • Pick specificity when choosing cards. The length of the personality field and the quality of the example dialogue predict session quality far better than the rating or download count.
  • If NPCs are what you care about, rank the directory by that axis rather than by overall score. The order is markedly different, and the per-axis winners guide names who takes each one.

Characters are the one problem this category has largely solved. The seven problems, ranked by how badly the field handles them covers the six it hasn’t.

How We Measured This

NPC Fidelity asks whether characters hold a consistent personality, pursue their own goals, and react durably to history — or whether they are agreeable mirrors. It is one of seven axes in methodology v1.0, scored 0–5 from first-hand play by Rukka, using a fixed probe set that includes social-state and consistency probes across both single- and multi-character configurations.

Two limits worth naming. Card platforms carry extra variance for the reasons above, and we publish that rather than hiding it. And we do not score pre-release software, so this is a reading of what shipped.

The figures here cover all 28 scored products, platforms and prompt-games together. The Arcanum AI RPG Benchmark publishes the platform half — the 16 scored platforms, every axis — where this axis averages higher still. If a number here doesn’t match your experience of the same product, we want to hear it: [email protected].

Frequently Asked Questions

Why do AI characters lose their personality over time? Usually one of two reasons. If the character was written vaguely, it never had a personality to hold — loose, impressionistic descriptions produce characters who shift to match whatever tone the conversation takes, which reads as responsiveness early and as having no identity later. If the character was written well, the cause is distance: as a session grows, the original description sits further from the model’s attention than the last twenty exchanges, so recent conversational tone gradually overwrites the defined personality.

Why do AI RPGs handle one character better than a group? Because holding one persona is a much smaller problem than holding several at once. With a single character the model has one voice to maintain and every line reinforces it. With a cast it must keep several distinct personalities separated while generating them from the same process, and the cheapest way to satisfy that is to converge them. Over a long session the distinct voices flatten into one narrator wearing several names — this is the single most common NPC failure we measured.

Which AI RPG has the best NPCs? Janitor AI, at 5 on NPC Fidelity — the only 5 we have awarded on that axis, to any product. Its characters hold personality across a long session and its multi-character cards keep their texture better than most. Much of that credit belongs to its card-writing community rather than its engineering, which is a real reason to choose it and worth stating plainly. Friends & Fables, Hidden Door and Character.AI all score 4.

What makes an AI NPC feel real? Specificity and independent goals. Specific detail gives the model something concrete to be consistent with, which is why a short, precise personality outperforms a long, atmospheric one. Independent goals are the rarer half: most AI NPCs react well and never act, so they respond convincingly to whatever you do and then wait. A character who wants something you have not offered them, and pursues it while you are elsewhere, is the thing almost nothing in this category does.

Do character cards matter more than the platform? On card platforms, a great deal. Where you supply a community-made card and the platform runs it, a strong card and a weak card can make the same platform feel like two different products. We hold the card selection constant across every platform in that class so the entries stay comparable to each other, which is why we are more confident in how those platforms rank against one another than in any single absolute number. On engine-driven platforms the system is the product and this problem does not arise.