Top

// research

The AI RPG Research Index

63 papers · 11 industry entries · last reviewed 2026-09-25 · download as JSON

The AI RPG Research Index is a dated, annotated list of academic work on the problems an AI RPG has to solve — remembering the story, keeping characters true, running the rules, and leaving room for the player — with each paper mapped to the axis of the Arcanum benchmark it bears on.

Research and products measure these problems in different ways. Nearly every paper here tests models, in controlled setups, against automated players or judges. Our benchmark tests platforms — the engine, the memory system, the rules and the model together — through first-hand play, because that is what a player actually uses. Neither can see what the other sees. Putting them side by side is the point of this page.

Every entry was checked against its primary source on 2026-09-25: the arXiv record for papers, the vendor's own page or the Steam store page for industry entries. Each finding is paraphrased from the paper itself. We do not score research, and nothing here is ranked.

What the Research Keeps Finding

Read together, these papers point the same way on four things. We look at each one in detail, including where the research and our own findings disagree, in what independent research says about AI RPGs.

  • Long stories break. In NCP-Bench, the best model kept its story intact in only 42% of runs after 20 turns. In RoleBreak, the strongest voice system broke character after about 10 turns on average. Our benchmark finds the same thing on shipped platforms: no platform has a top score on Memory & Continuity or Longevity.
  • Good writing does not mean a working game. RPGBench found models "can produce engaging stories but often struggle to implement consistent, verifiable game mechanics." NCP-Bench found that writing quality did not predict consistency.
  • Game state held outside the model helps. FIREBALL, function-calling game masters, PAYADOR and IVIE all improve consistency by giving the model a tracked world state instead of trusting the narration.
  • Open-ended AI characters are not automatically more fun. In a randomised study of 130 players, LLM-driven NPCs raised cognitive load and did not significantly improve the overall experience. Safety-aligned models also lose fidelity steadily as a character becomes less moral.

Memory and long-horizon consistency

Whether a system keeps its facts, commitments and characters straight as a story gets long — the problem our Memory & Continuity and Longevity axes measure.

PaperWhat it measures and foundBears on
RoleBreak: Benchmarking Long-Horizon Role-Playing Robustness in Spoken Dialogue Yuqi Wang et al. · 2026 Long-horizon role-play robustness in spoken (voice) dialogue: 310 roles, 6,688 human-verified turns. Even the strongest system hit its first persona failure after 10.4 turns on average; bigger models delayed failure but barely improved vocal emotion. What it can't see: Voice systems only; conversations, not games. NPC FidelityLongevity
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Yingpeng Ma et al. · 2026 · Accepted by ICML 2026 Whether an LLM narrator keeps its story commitments and established facts over long interactive narratives when a player pushes against them (NCP-Bench, 100 environments from movie synopses). The best model (GPT-5.2) kept its narrative intact in only 42% of runs after 20 turns; fact-conflict rates ran 40–68% across models, and high writing quality did not predict consistency. What it can't see: Tests models with an automated adversarial player, not shipped platforms with their own memory systems. Memory & ContinuityLongevityDeterminism & Fairness
Staying In Character: Perspective-Bounded Memory For Book-Based Role-Playing Agents Xushuo Tang et al. · 2026 Whether book-based character agents stay inside what their character could know (knowledge boundaries) and keep a varied voice. A three-layer memory (episodic, visibility-tagged facts, situational personality) improved knowledge-boundary fidelity by 34.6 points over the strongest prior method. What it can't see: Characters from novels, measured by question-answering, not play. Memory & ContinuityNPC Fidelity
DREAM: LLM-based Dynamic Role-playing via Event-Aware Memory Graph Zhihao Xiao et al. · 2026 · Accepted at KDD 2026 Temporal and causal consistency of role-playing agents (new TCM benchmark), using an event-aware memory graph (DREAM). Organising a character's experiences as a time-ordered, causally linked event graph beat strong baselines on CoSER, LifeChoice and TCM. What it can't see: Established literary characters; no player-driven campaign. Memory & ContinuityNPC Fidelity
From Facts to Insights: A Persona-Driven Dual Memory Framework and Dataset for Role-Playing Agents Rongsheng Zhang et al. · 2026 Whether memory systems interpret facts through the persona, not just store them (RoleMemo dataset, DualMem framework). Persona-agnostic summarisation yields generic answers; a 4B model with split factual/persona memory beat zero-shot DeepSeek-V3.2 frameworks on sustained persona fidelity. What it can't see: Preprint; dialogue tasks, not games. Memory & ContinuityNPC Fidelity
Memory-Driven Role-Playing: Evaluation and Enhancement of Persona Knowledge Utilization in LLMs Kai Wang et al. · 2026 Whether models recall and apply their persona knowledge from dialogue context alone (MREval: Anchoring, Recalling, Bounding, Enacting; MRBench). A structured retrieval prompt let small models (Qwen3-8B) match much larger closed models, and memory gains carried through to response quality. What it can't see: Single-character dialogue; 12 models, no platforms. Memory & ContinuityNPC Fidelity
Lost in Stories: Consistency Bugs in Long Story Generation by LLMs Junjie Li et al. · 2026 Consistency errors in long-form story generation (ConStory-Bench: 2,000 prompts, 19 error subtypes). Contradictions cluster in factual and temporal details and tend to appear around the middle of long narratives. What it can't see: Non-interactive story generation. Memory & ContinuityLongevity
Fixed-Persona SLMs with Modular Memory: Scalable NPC Dialogue on Consumer Hardware Martin Braas and Lukas Esterle · 2025 NPC dialogue with small fine-tuned models plus swappable memory modules on consumer hardware. Persona-tuned small models with runtime-swappable memory kept character context without reloading, on consumer GPUs. What it can't see: Engineering study with small models; no player evaluation of fun. Memory & ContinuityNPC Fidelity
MOOM: Maintenance, Organization and Optimization of Memory in Ultra-Long Role-Playing Dialogues Weishu Chen et al. · 2025 Memory extraction for ultra-long role-play dialogues (Chinese ZH-4O dataset, ~600 turns per dialogue). A dual-branch memory (plot conflicts + user profile) with a forgetting mechanism kept memory size controlled while beating prior extraction methods. What it can't see: Chinese-language dialogue data; memory extraction measured, not play experience. Memory & ContinuityLongevity
StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns Luanbo Wan and Weizhi Ma · 2025 Long-term memory measured through branching interactive-fiction games (StoryBench). Interactive fiction works as a test bed for long-term memory: choices cascade across turns, and models must trace back and revise earlier decisions. What it can't see: Measures models as players, not as narrators. Memory & ContinuityPlayer Agency
RecurrentGPT: Interactive Generation of (Arbitrarily) Long Text Wangchunshu Zhou et al. · 2023 Generating arbitrarily long text by simulating long/short-term memory in natural language (RecurrentGPT). Human-readable, editable memories let an LLM write past its context window; demonstrated as personalised interactive fiction. What it can't see: 2023 system; predates today's long-context models. Memory & ContinuityLongevity

Character and NPC fidelity

Whether characters stay themselves: their voice, their knowledge, their morality, and how players experience them.

PaperWhat it measures and foundBears on
CHARM: Character Hallucination for Multicultural Role Play Benchmark Sunkyung Han et al. · 2026 · Accepted to Findings of EMNLP 2026 Whether characters respect knowledge boundaries across five cultural regions (CHARM, 40 characters). Hallucination comes mostly from compliance failures: models recognise a question is out of character and answer it anyway. What it can't see: Multiple-choice probes. NPC Fidelity
ArcANE: Do Role-Playing Language Agents Stay in Character at the Right Time? Woojung Song et al. · 2026 · Accepted at EMNLP 2026 (Main) Whether role-playing agents follow a character's development over a story rather than a fixed persona (ArcANE). Giving the model the character's arc up to the current chapter beat every other context strategy by 2.2–8.4 points in all six models. What it can't see: Novel characters, question-style scenarios. NPC FidelityMemory & Continuity
The Double-Edged Sword of Open-Ended Interaction: How LLM-Driven NPCs Affect Players' Cognitive Load and Gaming Experience Ting-Chen Hsu et al. · 2026 How LLM NPCs affect players' cognitive load and experience versus scripted NPCs (randomised study, N=130). LLM NPCs significantly raised cognitive load and did not significantly improve overall experience; autonomy rose while usability and trust fell. What it can't see: One research prototype game. Player AgencyNPC FidelitySignature Design
Too Good to be Bad: On the Failure of LLMs to Role-Play Villains Zihao Yi et al. · 2025 How faithfully models play characters across a four-level moral scale, from paragons to villains (Moral RolePlay). Fidelity falls steadily as characters become less moral; safety-aligned models swap nuanced malice for surface aggression, and chatbot skill does not predict villain skill. What it can't see: Scripted scenes; measures the model, not a platform's prompt design. NPC FidelitySignature Design
Symbolically Scaffolded Play: Designing Role-Sensitive Prompts for Generative NPC Dialogue Vanessa Figueiredo and David Elumeze · 2025 Whether tighter prompt constraints improve player experience with generative NPCs (voice detective game, N=10 study + LLM judge). Scaffolding helped the quest-giver NPC but made suspect NPCs less believable — tighter constraints do not automatically make play better. What it can't see: Very small user study. NPC FidelityPlayer Agency
Deflanderization for Game Dialogue: Balancing Character Authenticity with Task Execution in LLM-based NPCs Pasin Buakhaw et al. · 2025 NPCs that must both stay in character and complete game tasks (Commonsense Persona-Grounded Dialogue Challenge 2025). Prompting that suppresses excessive role-play ("Deflanderization") improved task fidelity; competition entry placed 2nd on two tasks. What it can't see: Competition report, short. NPC FidelityMechanical Depth
Tricking LLM-Based NPCs into Spilling Secrets Kyohei Shiomi et al. · 2025 Whether prompt injection can make LLM NPCs reveal secrets they are meant to keep. Adversarial prompts can make LLM-based NPCs spill hidden background secrets — a security issue for game design. What it can't see: Short study. NPC FidelityDeterminism & Fairness
FAIRGAMER: Evaluating Social Biases in LLM-Based Video Game NPCs Bingkang Shi et al. · 2025 Social bias in LLM NPC decisions across trade, cooperation and competition (FairGamer). All seven frontier models showed biased decisions, and larger models showed more bias. What it can't see: Decision tasks, not dialogue quality. NPC FidelityDeterminism & Fairness
RMTBench: Benchmarking LLMs Through Multi-Turn User-Centric Role-Playing Hao Xiang et al. · 2025 User-centric multi-turn role-play (RMTBench, 80 characters, 8,000+ rounds). Built dialogues around what the user wants rather than the character sheet, closer to real use. What it can't see: Simulated users, LLM scoring. NPC FidelityPlayer Agency
An Empirical Evaluation of AI-Powered Non-Player Characters' Perceived Realism and Performance in Virtual Reality Environments Mikko Korkiakoski et al. · 2025 Perceived realism and latency of GPT-4-driven NPCs in a VR interrogation game (N=18). Believability 6.67/10 and a 7-second average response cycle that grew with conversation context. What it can't see: Small sample; VR voice setting. NPC Fidelity
Guess What I am Thinking: A Benchmark for Inner Thought Reasoning of Role-Playing Language Agents Rui Xu et al. · 2025 Whether role-playing agents can generate a character's inner thoughts (RoleThink). Retrieving memories and predicting reactions before answering improved inner-thought generation. What it can't see: Literary characters. NPC Fidelity
CoSER: A Comprehensive Literary Dataset and Framework for Training and Evaluating LLM Role-Playing and Persona Simulation Xintao Wang et al. · 2025 · Accepted by ICML 2025 Training and evaluating role-play of established characters from 771 books (CoSER, 17,966 characters). Its open 70B model matched or beat GPT-4o on its own and three existing benchmarks. What it can't see: Book scenes, not games. NPC Fidelity
CharacterBench: Benchmarking Character Customization of Large Language Models Jinfeng Zhou et al. · 2024 · AAAI 2025 Character customisation ability of LLMs across 11 dimensions (CharacterBench, 22,859 human-annotated samples, 3,956 characters). Largest bilingual character benchmark; its trained judge beat GPT-4 as an evaluator. What it can't see: Dialogue snapshots, not sustained play. NPC Fidelity
CharacterBox: Evaluating the Role-Playing Capabilities of LLMs in Text-Based Virtual Worlds Lei Wang et al. · 2024 Role-play in simulated text worlds with a character agent and a narrator agent (CharacterBox). Behaviour trajectories in a sandbox give a deeper read of role-play than Q&A snapshots. What it can't see: Simulated users; no human players. NPC FidelityPlayer Agency
PingPong: A Benchmark for Role-Playing Language Models with User Emulation and Multi-Model Evaluation Ilya Gusev · 2024 Role-play quality via simulated users and a judge ensemble: character consistency, entertainment, fluency (PingPong, 40+ models). Automated judgments correlated strongly with human annotations. What it can't see: LLM-judged; English and Russian. NPC FidelitySignature Design
TimeChara: Evaluating Point-in-Time Character Hallucination of Role-Playing Large Language Models Jaewoo Ahn et al. · 2024 · ACL 2024 Findings Point-in-time character hallucination: characters knowing things they should not yet know (TimeChara, 10,895 instances). Significant hallucination even in GPT-4o. What it can't see: Fandom characters, question format. NPC FidelityMemory & Continuity
CharacterEval: A Chinese Benchmark for Role-Playing Conversational Agent Evaluation Quan Tu et al. · 2024 Chinese role-playing conversational agents across 13 metrics (CharacterEval, 77 characters). Chinese LLMs outperformed GPT-4 in Chinese role-play conversation. What it can't see: Chinese-language only. NPC Fidelity
InCharacter: Evaluating Personality Fidelity in Role-Playing Agents through Psychological Interviews Xintao Wang et al. · 2023 · ACL 2024 Personality fidelity of role-playing agents measured by psychological interviews (InCharacter, 32 characters, 14 scales). State-of-the-art agents matched human-perceived character personalities with up to 80.7% accuracy. What it can't see: Personality tests, not gameplay. NPC Fidelity
RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models Zekun Moore Wang et al. · 2023 Benchmarking and improving role-playing ability (RoleBench, 168,093 samples, 100 roles). Role-conditioned tuning let open models approach GPT-4 role-play. What it can't see: 2023 models; speaking-style focus. NPC Fidelity

Game state, rules and the AI game master

Whether an AI can run a game rather than tell a story — tracking state, enforcing rules, and assisting or replacing a game master.

PaperWhat it measures and foundBears on
Orchestrated Reality: From Role-Play to Living, Playable Game Worlds -- LLM-Driven World Simulation as a Parameterized-Action POMDP Yuhang Huang et al. · 2026 Architecture for LLM-run game worlds where the world is a canonical object owned by an orchestrator (work in progress). Argues that deployed systems let the narrator assert state in free prose "without any validated representation", so a fully autonomous engine remains infeasible. What it can't see: Work in progress; no evaluation yet. Mechanical DepthMemory & ContinuityDeterminism & Fairness
IVIE: A Neuro-symbolic Approach to Incremental and Validated Generation of Interactive Fiction Worlds Micaela Vaucher et al. · 2026 · To appear in the Proceedings of the 16th International Conference on Computational Creativity (ICCC'26), June 2026 Generating complete, playable interactive-fiction worlds with symbolic validation (IVIE). Symbolic validation grounded the LLM without removing creative freedom, though model inconsistencies occasionally bypassed puzzles. What it can't see: Generation of IF worlds, not live campaigns. Mechanical DepthDeterminism & Fairness
World-State Transformations for Neuro-symbolic Interactive Storytelling Santiago Góngora et al. · 2026 · To be presented at the 17th International Conference on Computational Creativity (ICCC'26) Neuro-symbolic interactive storytelling where the LLM triggers pre-programmed world-state changes. World-state transformations kept the world consistent while still encouraging creative player input. What it can't see: Eight participants; exploratory. Mechanical DepthDeterminism & FairnessPlayer Agency
From World-Gen to Quest-Line: A Dependency-Driven Prompt Pipeline for Coherent RPG Generation Dominik Borawski et al. · 2026 A staged, schema-enforced prompt pipeline for generating RPG worlds, NPCs and quests. Structured JSON hand-offs between stages reduced drift and kept content valid as complexity grew. What it can't see: Qualitative evaluation. Mechanical DepthSignature Design
First Steps Towards Overhearing LLM Agents: A Case Study With Dungeons & Dragons Gameplay Andrew Zhu et al. · 2025 · COLM 2025 Workshop on AI Agents "Overhearing" agents that listen to a D&D table and help the DM in the background. Some large audio-language models can perform background assistance from implicit audio cues. What it can't see: Workshop paper; assistive, not autonomous. Mechanical Depth
STORY2GAME: Generating (Almost) Everything in an Interactive Fiction Game Eric Zhou et al. · 2025 Generating interactive-fiction games including the action code, with new actions created on demand (STORY2GAME). Generating action preconditions and effects lets stories stay open-ended yet grounded in tracked game state. What it can't see: Generated games, not live play. Mechanical DepthPlayer Agency
PAYADOR: A Minimalist Approach to Grounding Language Models on Structured Data for Interactive Storytelling and Role-playing Games Santiago Góngora et al. · 2025 · Proceedings of the Fifteenth International Conference on Computational Creativity Grounding an LLM on a minimal world representation by predicting action outcomes (PAYADOR). Predicting outcomes instead of mapping input to fixed actions preserves player freedom in RPG-style play. What it can't see: Minimal prototype. Mechanical DepthDeterminism & FairnessPlayer Agency
RPGBENCH: Evaluating Large Language Models as Role-Playing Game Engines Pengfei Yu et al. · 2025 LLMs as text RPG engines: creating a valid game (event-state representation) and simulating play while updating state and enforcing rules (RPGBench). State-of-the-art LLMs produce engaging stories but often fail to implement consistent, verifiable game mechanics, especially in long or complex scenarios. What it can't see: Models tested as engines, not shipped platforms with external state. Mechanical DepthDeterminism & FairnessMemory & Continuity
You Have Thirteen Hours in Which to Solve the Labyrinth: Enhancing AI Game Masters with Function Calling Jaewoo Song et al. · 2024 · Wordplay Workshop @ ACL 2024 AI game masters using function calling for game-specific controls in a tabletop RPG (Labyrinth: The Adventure Game). Function calling improved both narrative quality and state-update consistency, by human evaluation and unit tests. What it can't see: One game system; workshop paper. Mechanical DepthDeterminism & FairnessMemory & Continuity
LARP: Language-Agent Role Play for Open-World Games Ming Yan et al. · 2023 Cognitive architecture for role-playing agents in open-world games (LARP). Combines memory processing, learnable action space and personality alignment for open-world NPCs. What it can't see: Framework paper. Memory & ContinuityNPC Fidelity
CALYPSO: LLMs as Dungeon Masters' Assistants Andrew Zhu et al. · 2023 · AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment (AIIDE) 2023 LLM tools that assist a human Dungeon Master (CALYPSO). DMs found the output good enough to read to players and used it for ideas while keeping creative control. What it can't see: Assists a human DM rather than replacing one. Mechanical DepthSignature Design
FIREBALL: A Dataset of Dungeons and Dragons Actual-Play with Structured Game State Information Andrew Zhu et al. · 2023 · Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics D&D actual play with true game state from ~25,000 Discord sessions using the Avrae bot (FIREBALL). Giving models real game-state information improved generated turns; models can also produce executable game commands. What it can't see: Human play logs; model generation evaluated, not full GM play. Mechanical DepthDeterminism & Fairness
Rolling the Dice: Imagining Generative AI as a Dungeons & Dragons Storytelling Companion Jose Ma. Santiago et al. · 2023 Design vision for generative AI as a D&D storytelling companion. Proposes design guidelines and flags immersion and cognitive-load questions. What it can't see: Position paper, no experiment. Player AgencySignature Design
Dungeons and Dragons as a Dialog Challenge for Artificial Intelligence Chris Callison-Burch et al. · 2022 · Conference on Empirical Methods in Natural Language Processing (EMNLP), pp D&D as a dialogue challenge: generating the next turn and predicting game state (~900 games, 800,000 turns). Tracking game state measurably improves the model's turns; frames D&D as a benchmark problem for AI. What it can't see: Pre-ChatGPT models. Mechanical DepthNPC Fidelity

Player agency and narrative control

How much a player can change the story, and how systems keep authored structure without taking that away.

PaperWhat it measures and foundBears on
Design Techniques for LLM-Powered Interactive Storytelling: A Case Study of the Dramamancer System Tiffany Wang et al. · 2026 · Wordplay Workshop at EMNLP Design techniques for turning author-written story schemas into player-driven play (Dramamancer). Frames the author-intent vs player-agency balance as a design problem. What it can't see: Extended abstract. Player AgencySignature Design
SNAP: A Plan-Driven Framework for Controllable Interactive Narrative Generation Geonwoo Bang et al. · 2025 Plan-driven control of interactive narratives to prevent drift (SNAP). Splitting the story into planned cells kept dialogue scenario-consistent under varied user input. What it can't see: Short paper. Player AgencyMemory & Continuity
WHAT-IF: Exploring Branching Narratives by Meta-Prompting Large Language Models Runsheng "Anson" Huang et al. · 2024 · Published in Wordplay: When Language Meets Games Workshop (EMNLP 2025) Branching interactive fiction generated from a linear story by meta-prompting (WHAT-IF). Storing the branch tree as a graph kept alternate storylines coherent. What it can't see: Choice-based IF, not free text. Player Agency
Player-Driven Emergence in LLM-Driven Game Narrative Xiangyu Peng et al. · 2024 · IEEE Conference on Games 2024 Emergent narrative from player interaction with GPT-4 NPCs in a mystery text adventure (28 players). Players discovered new story nodes not in the original narrative; the most "emergent" players were those who like exploration games. What it can't see: One game, small sample. Player AgencySignature Design
GENEVA: GENErating and Visualizing branching narratives using LLMs Jorge Leandro et al. · 2023 · Accepted at IEEE Conference on Games 2024 LLM tool that generates branching and reconverging narrative graphs for designers (GENEVA). GPT-4 produced rich branching narratives under designer constraints. What it can't see: Authoring tool, not a runtime GM. Player AgencySignature Design

AI that plays games (context)

Benchmarks where the AI is the player, not the game master. Useful context, but they measure a different problem from ours.

PaperWhat it measures and foundBears on
ByteSized32Refactored: Towards an Extensible Interactive Text Games Corpus for LLM World Modeling and Evaluation Haonan Wang et al. · 2025 · Accepted to the 5th Wordplay: When Language Meets Games Workshop, EMNLP 2025 Corpus of 32 text games for LLM world modelling (ByteSized32Refactored). Refactored, extensible text-game corpus for generation and evaluation. What it can't see: Resource paper. —
Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games Dongmin Park et al. · 2025 LLM agents playing 12 popular video games across genres (Orak). Plug-and-play benchmark and fine-tuning data for game-playing agents. What it can't see: Measures AI as the player, not as the game master. —
TALES: Text Adventure Learning Environment Suite Christopher Zhang Cui et al. · 2025 LLMs playing synthetic and human-written text adventures (TALES). Even top agents score under 15% on text adventures designed for human enjoyment. What it can't see: AI as player. Memory & Continuity
TextArena Leon Guertler et al. · 2025 Competitive text games for agent evaluation (TextArena, 57+ environments). Online leaderboard with TrueSkill for social skills like negotiation and deception. What it can't see: AI as player. —
DSGBench: A Diverse Strategic Game Benchmark for Evaluating LLM-based Agents in Complex Decision-Making Environments Wenjie Tang et al. · 2025 Strategic decision-making in six complex games (DSGBench). Fine-grained scoring of agents in long-horizon strategy games. What it can't see: AI as player. —
Can VLMs Play Action Role-Playing Games? Take Black Myth Wukong as a Study Case Peng Chen et al. · 2024 Vision-language agents playing Black Myth: Wukong. Explores playing an action RPG from screen input only. What it can't see: Combat play, not narrative. —
Towards a Holodeck-style Simulation Game Ahad Shams et al. · 2023 A Holodeck-style simulation game built on generative agents (Infinitia). Applies the Generative Agents idea to a playable, multiplayer Unity sandbox. What it can't see: System description. Signature Design
Generative Agents: Interactive Simulacra of Human Behavior Joon Sung Park et al. · 2023 Believable agents with memory, reflection and planning in a Sims-like town (Generative Agents). The foundational architecture — store experiences, reflect, retrieve to plan — behind most "AI NPC" work since. What it can't see: Social simulation, not RPG play. NPC FidelityMemory & Continuity
Can Large Language Models Play Text Games Well? Current State-of-the-Art and Open Questions Chen Feng Tsai et al. · 2023 How well ChatGPT plays text games (2023). ChatGPT could not build a world model from play or even the manual. What it can't see: Early model generation. Memory & Continuity

Surveys and field maps

Reviews that map the role-playing agent literature as a whole.

PaperWhat it measures and foundBears on
Role-Playing Agents Driven by Large Language Models: Current Status, Challenges, and Future Trends Ye Wang et al. · 2026 2026 review of role-playing agents: personality modelling, memory, evaluation. Traces the field from templates to cognitive simulation. What it can't see: Review. NPC FidelityMemory & Continuity
Towards a Design Guideline for RPA Evaluation: A Survey of Large Language Model-Based Role-Playing Agents Chaoran Chen et al. · 2025 Evaluation-design guideline from a review of 1,676 role-playing agent papers. Identifies six agent attributes, seven task attributes and seven evaluation metrics used across the literature. What it can't see: Guideline for researchers. NPC Fidelity
The Oscars of AI Theater: A Survey on Role-Playing with Language Models Nuo Chen et al. · 2024 Survey of role-playing with language models ("The Oscars of AI Theater"). Taxonomy of data, models, architecture and evaluation for role-play. What it can't see: Review. NPC Fidelity
From Persona to Personalization: A Survey on Role-Playing Language Agents Jiangjie Chen et al. · 2024 · Accepted to TMLR 2024 Survey of role-playing language agents: demographic, character and individualised personas. The most-cited field map of role-playing agents. What it can't see: General role-play, games are one application. NPC FidelityMemory & Continuity
A Survey on the Memory Mechanism of Large Language Model based Agents Zeyu Zhang et al. · 2024 Survey of memory mechanisms in LLM-based agents: design and evaluation of memory modules. The standard reference for how agent memory is built and evaluated. What it can't see: General agents, not games. Memory & Continuity

In Industry: Where These Ideas Ship

The research above shows up in shipped products — middleware for game studios, mods, and games on Steam that run language models live. These entries are descriptive and unrated: we have not played them for the benchmark, so we say what each one is, in its maker's own words, and nothing about how good it is. For a player's guide to the games, see games with AI NPCs you can talk to; for NVIDIA's character toolkit, NVIDIA ACE explained; for the two best-known character platforms, Convai and what happened to Inworld AI.

NameStatusWhat it is
NVIDIA ACENVIDIA shipping (SDK) NVIDIA describes ACE as "a suite of digital human technologies that power agentic workflows for autonomous game characters and digital assistants." Source: vendor developer page
Inworld AIInworld pivoted Once known for game-character AI; its homepage now sells "Realtime TTS and STT models, LLM serving, and the inference behind both" for consumer-facing applications. Source: vendor homepage
ConvaiConvai shipping Positions itself as "Conversational AI for Virtual Worlds." Source: vendor homepage title
Ubisoft NEO NPC (R&D)Ubisoft research prototype A small R&D team at Ubisoft Paris experimenting with generative AI toward real conversations with NPCs, per Ubisoft's own news post. Source: Ubisoft News
Mantellaart-from-the-machine (open source) active (repo pushed 2026-07-16) A Skyrim and Fallout 4 mod "which allows you to naturally speak to NPCs using a Speech-to-Text → LLMs → Text-to-Speech pipeline." Source: GitHub repository
AI RogueliteSteam app 1889620 released 2023-10-25 Developer disclosure on Steam: it "heavily uses AI to live-generate in-game content such as text, images, and sound effects" and "to make a variety of game mechanics decisions in real time." Source: Steam store page AI disclosure
inZOISteam app 2456740 (Krafton) early access since 2025-03-27 Steam disclosure: player text can influence "character actions and thoughts … using SLM technology", plus AI-generated textures, 3D objects and motions. Source: Steam store page AI disclosure
Where Winds MeetSteam app 3564740 released 2025-11-14 Steam disclosure: "AI-driven NPC text and voice chat that responds to player input in real time." Source: Steam store page AI disclosure
Suck Up!Steam app 2726370 (Proxima) released 2025-10-01 Steam disclosure: players talk to AI characters by voice and "the AI responds in real time based on tone and strategy." Source: Steam store page AI disclosure
WanderfolkSteam app 4599270 (Kinetic Sky) coming 2026 Steam disclosure: villager conversations "generated in real time by a large language model", and NPCs "remember what you've said to them across the playthrough." Source: Steam store page AI disclosure
VaudevilleSteam app 2240920 (Bumblebee Studios) released 2025-11-28 Steam disclosure: "dialogues generated with the help of a conversational AI"; players talk to townsfolk by typing or voice. Source: Steam store page AI disclosure

How This Index Is Built

  • What gets in. Research that bears on at least one of our seven axes, or on how an AI RPG is built. We searched arXiv for role-playing, game-master, text-adventure, interactive-fiction and NPC work, read every candidate's abstract, and kept 63 of about 255. Papers where the AI is the player are kept as context, in their own group.
  • How it is checked. Titles, authors and dates come from the arXiv record, not from memory or from another summary. A script re-checks every title against arXiv and every industry link against its source before each update.
  • Venues. We show a conference or journal only when the authors state it on arXiv. A paper that says it was submitted somewhere is listed as a preprint, because submitted is not accepted.
  • Freshness. The index is re-reviewed every quarter, alongside each benchmark deposit. New papers that bear on the axes are added; nothing is quietly removed.
  • Corrections. If we have misread a paper, tell us. Mistakes we fix are logged on the corrections page, like any other.

Take the Data

The whole index is available as research.json under CC BY 4.0: every entry, its source, its axis mapping and the date it was verified. Our summaries and mappings are free to reuse with attribution. The papers belong to their authors — we link to them rather than reproduce them.

For the measurements this index is set against, see the benchmark. For how those measurements are made, see the methodology.