// research
The AI RPG Research Index
63 papers · 11 industry entries · last reviewed 2026-09-25 · download as JSON
The AI RPG Research Index is a dated, annotated list of academic work on the problems an AI RPG has to solve — remembering the story, keeping characters true, running the rules, and leaving room for the player — with each paper mapped to the axis of the Arcanum benchmark it bears on.
Research and products measure these problems in different ways. Nearly every paper here tests models, in controlled setups, against automated players or judges. Our benchmark tests platforms — the engine, the memory system, the rules and the model together — through first-hand play, because that is what a player actually uses. Neither can see what the other sees. Putting them side by side is the point of this page.
Every entry was checked against its primary source on 2026-09-25: the arXiv record for papers, the vendor's own page or the Steam store page for industry entries. Each finding is paraphrased from the paper itself. We do not score research, and nothing here is ranked.
What the Research Keeps Finding
Read together, these papers point the same way on four things. We look at each one in detail, including where the research and our own findings disagree, in what independent research says about AI RPGs.
- Long stories break. In NCP-Bench, the best model kept its story intact in only 42% of runs after 20 turns. In RoleBreak, the strongest voice system broke character after about 10 turns on average. Our benchmark finds the same thing on shipped platforms: no platform has a top score on Memory & Continuity or Longevity.
- Good writing does not mean a working game. RPGBench found models "can produce engaging stories but often struggle to implement consistent, verifiable game mechanics." NCP-Bench found that writing quality did not predict consistency.
- Game state held outside the model helps. FIREBALL, function-calling game masters, PAYADOR and IVIE all improve consistency by giving the model a tracked world state instead of trusting the narration.
- Open-ended AI characters are not automatically more fun. In a randomised study of 130 players, LLM-driven NPCs raised cognitive load and did not significantly improve the overall experience. Safety-aligned models also lose fidelity steadily as a character becomes less moral.
Memory and long-horizon consistency
Whether a system keeps its facts, commitments and characters straight as a story gets long — the problem our Memory & Continuity and Longevity axes measure.
| Paper | What it measures and found | Bears on |
|---|---|---|
| RoleBreak: Benchmarking Long-Horizon Role-Playing Robustness in Spoken Dialogue | Long-horizon role-play robustness in spoken (voice) dialogue: 310 roles, 6,688 human-verified turns. Even the strongest system hit its first persona failure after 10.4 turns on average; bigger models delayed failure but barely improved vocal emotion. What it can't see: Voice systems only; conversations, not games. | NPC FidelityLongevity |
| Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives | Whether an LLM narrator keeps its story commitments and established facts over long interactive narratives when a player pushes against them (NCP-Bench, 100 environments from movie synopses). The best model (GPT-5.2) kept its narrative intact in only 42% of runs after 20 turns; fact-conflict rates ran 40–68% across models, and high writing quality did not predict consistency. What it can't see: Tests models with an automated adversarial player, not shipped platforms with their own memory systems. | Memory & ContinuityLongevityDeterminism & Fairness |
| Staying In Character: Perspective-Bounded Memory For Book-Based Role-Playing Agents | Whether book-based character agents stay inside what their character could know (knowledge boundaries) and keep a varied voice. A three-layer memory (episodic, visibility-tagged facts, situational personality) improved knowledge-boundary fidelity by 34.6 points over the strongest prior method. What it can't see: Characters from novels, measured by question-answering, not play. | Memory & ContinuityNPC Fidelity |
| DREAM: LLM-based Dynamic Role-playing via Event-Aware Memory Graph | Temporal and causal consistency of role-playing agents (new TCM benchmark), using an event-aware memory graph (DREAM). Organising a character's experiences as a time-ordered, causally linked event graph beat strong baselines on CoSER, LifeChoice and TCM. What it can't see: Established literary characters; no player-driven campaign. | Memory & ContinuityNPC Fidelity |
| From Facts to Insights: A Persona-Driven Dual Memory Framework and Dataset for Role-Playing Agents | Whether memory systems interpret facts through the persona, not just store them (RoleMemo dataset, DualMem framework). Persona-agnostic summarisation yields generic answers; a 4B model with split factual/persona memory beat zero-shot DeepSeek-V3.2 frameworks on sustained persona fidelity. What it can't see: Preprint; dialogue tasks, not games. | Memory & ContinuityNPC Fidelity |
| Memory-Driven Role-Playing: Evaluation and Enhancement of Persona Knowledge Utilization in LLMs | Whether models recall and apply their persona knowledge from dialogue context alone (MREval: Anchoring, Recalling, Bounding, Enacting; MRBench). A structured retrieval prompt let small models (Qwen3-8B) match much larger closed models, and memory gains carried through to response quality. What it can't see: Single-character dialogue; 12 models, no platforms. | Memory & ContinuityNPC Fidelity |
| Lost in Stories: Consistency Bugs in Long Story Generation by LLMs | Consistency errors in long-form story generation (ConStory-Bench: 2,000 prompts, 19 error subtypes). Contradictions cluster in factual and temporal details and tend to appear around the middle of long narratives. What it can't see: Non-interactive story generation. | Memory & ContinuityLongevity |
| Fixed-Persona SLMs with Modular Memory: Scalable NPC Dialogue on Consumer Hardware | NPC dialogue with small fine-tuned models plus swappable memory modules on consumer hardware. Persona-tuned small models with runtime-swappable memory kept character context without reloading, on consumer GPUs. What it can't see: Engineering study with small models; no player evaluation of fun. | Memory & ContinuityNPC Fidelity |
| MOOM: Maintenance, Organization and Optimization of Memory in Ultra-Long Role-Playing Dialogues | Memory extraction for ultra-long role-play dialogues (Chinese ZH-4O dataset, ~600 turns per dialogue). A dual-branch memory (plot conflicts + user profile) with a forgetting mechanism kept memory size controlled while beating prior extraction methods. What it can't see: Chinese-language dialogue data; memory extraction measured, not play experience. | Memory & ContinuityLongevity |
| StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns | Long-term memory measured through branching interactive-fiction games (StoryBench). Interactive fiction works as a test bed for long-term memory: choices cascade across turns, and models must trace back and revise earlier decisions. What it can't see: Measures models as players, not as narrators. | Memory & ContinuityPlayer Agency |
| RecurrentGPT: Interactive Generation of (Arbitrarily) Long Text | Generating arbitrarily long text by simulating long/short-term memory in natural language (RecurrentGPT). Human-readable, editable memories let an LLM write past its context window; demonstrated as personalised interactive fiction. What it can't see: 2023 system; predates today's long-context models. | Memory & ContinuityLongevity |
Character and NPC fidelity
Whether characters stay themselves: their voice, their knowledge, their morality, and how players experience them.
| Paper | What it measures and found | Bears on |
|---|---|---|
| CHARM: Character Hallucination for Multicultural Role Play Benchmark | Whether characters respect knowledge boundaries across five cultural regions (CHARM, 40 characters). Hallucination comes mostly from compliance failures: models recognise a question is out of character and answer it anyway. What it can't see: Multiple-choice probes. | NPC Fidelity |
| ArcANE: Do Role-Playing Language Agents Stay in Character at the Right Time? | Whether role-playing agents follow a character's development over a story rather than a fixed persona (ArcANE). Giving the model the character's arc up to the current chapter beat every other context strategy by 2.2–8.4 points in all six models. What it can't see: Novel characters, question-style scenarios. | NPC FidelityMemory & Continuity |
| The Double-Edged Sword of Open-Ended Interaction: How LLM-Driven NPCs Affect Players' Cognitive Load and Gaming Experience | How LLM NPCs affect players' cognitive load and experience versus scripted NPCs (randomised study, N=130). LLM NPCs significantly raised cognitive load and did not significantly improve overall experience; autonomy rose while usability and trust fell. What it can't see: One research prototype game. | Player AgencyNPC FidelitySignature Design |
| Too Good to be Bad: On the Failure of LLMs to Role-Play Villains | How faithfully models play characters across a four-level moral scale, from paragons to villains (Moral RolePlay). Fidelity falls steadily as characters become less moral; safety-aligned models swap nuanced malice for surface aggression, and chatbot skill does not predict villain skill. What it can't see: Scripted scenes; measures the model, not a platform's prompt design. | NPC FidelitySignature Design |
| Symbolically Scaffolded Play: Designing Role-Sensitive Prompts for Generative NPC Dialogue | Whether tighter prompt constraints improve player experience with generative NPCs (voice detective game, N=10 study + LLM judge). Scaffolding helped the quest-giver NPC but made suspect NPCs less believable — tighter constraints do not automatically make play better. What it can't see: Very small user study. | NPC FidelityPlayer Agency |
| Deflanderization for Game Dialogue: Balancing Character Authenticity with Task Execution in LLM-based NPCs | NPCs that must both stay in character and complete game tasks (Commonsense Persona-Grounded Dialogue Challenge 2025). Prompting that suppresses excessive role-play ("Deflanderization") improved task fidelity; competition entry placed 2nd on two tasks. What it can't see: Competition report, short. | NPC FidelityMechanical Depth |
| Tricking LLM-Based NPCs into Spilling Secrets | Whether prompt injection can make LLM NPCs reveal secrets they are meant to keep. Adversarial prompts can make LLM-based NPCs spill hidden background secrets — a security issue for game design. What it can't see: Short study. | NPC FidelityDeterminism & Fairness |
| FAIRGAMER: Evaluating Social Biases in LLM-Based Video Game NPCs | Social bias in LLM NPC decisions across trade, cooperation and competition (FairGamer). All seven frontier models showed biased decisions, and larger models showed more bias. What it can't see: Decision tasks, not dialogue quality. | NPC FidelityDeterminism & Fairness |
| RMTBench: Benchmarking LLMs Through Multi-Turn User-Centric Role-Playing | User-centric multi-turn role-play (RMTBench, 80 characters, 8,000+ rounds). Built dialogues around what the user wants rather than the character sheet, closer to real use. What it can't see: Simulated users, LLM scoring. | NPC FidelityPlayer Agency |
| An Empirical Evaluation of AI-Powered Non-Player Characters' Perceived Realism and Performance in Virtual Reality Environments | Perceived realism and latency of GPT-4-driven NPCs in a VR interrogation game (N=18). Believability 6.67/10 and a 7-second average response cycle that grew with conversation context. What it can't see: Small sample; VR voice setting. | NPC Fidelity |
| Guess What I am Thinking: A Benchmark for Inner Thought Reasoning of Role-Playing Language Agents | Whether role-playing agents can generate a character's inner thoughts (RoleThink). Retrieving memories and predicting reactions before answering improved inner-thought generation. What it can't see: Literary characters. | NPC Fidelity |
| CoSER: A Comprehensive Literary Dataset and Framework for Training and Evaluating LLM Role-Playing and Persona Simulation | Training and evaluating role-play of established characters from 771 books (CoSER, 17,966 characters). Its open 70B model matched or beat GPT-4o on its own and three existing benchmarks. What it can't see: Book scenes, not games. | NPC Fidelity |
| CharacterBench: Benchmarking Character Customization of Large Language Models | Character customisation ability of LLMs across 11 dimensions (CharacterBench, 22,859 human-annotated samples, 3,956 characters). Largest bilingual character benchmark; its trained judge beat GPT-4 as an evaluator. What it can't see: Dialogue snapshots, not sustained play. | NPC Fidelity |
| CharacterBox: Evaluating the Role-Playing Capabilities of LLMs in Text-Based Virtual Worlds | Role-play in simulated text worlds with a character agent and a narrator agent (CharacterBox). Behaviour trajectories in a sandbox give a deeper read of role-play than Q&A snapshots. What it can't see: Simulated users; no human players. | NPC FidelityPlayer Agency |
| PingPong: A Benchmark for Role-Playing Language Models with User Emulation and Multi-Model Evaluation | Role-play quality via simulated users and a judge ensemble: character consistency, entertainment, fluency (PingPong, 40+ models). Automated judgments correlated strongly with human annotations. What it can't see: LLM-judged; English and Russian. | NPC FidelitySignature Design |
| TimeChara: Evaluating Point-in-Time Character Hallucination of Role-Playing Large Language Models | Point-in-time character hallucination: characters knowing things they should not yet know (TimeChara, 10,895 instances). Significant hallucination even in GPT-4o. What it can't see: Fandom characters, question format. | NPC FidelityMemory & Continuity |
| CharacterEval: A Chinese Benchmark for Role-Playing Conversational Agent Evaluation | Chinese role-playing conversational agents across 13 metrics (CharacterEval, 77 characters). Chinese LLMs outperformed GPT-4 in Chinese role-play conversation. What it can't see: Chinese-language only. | NPC Fidelity |
| InCharacter: Evaluating Personality Fidelity in Role-Playing Agents through Psychological Interviews | Personality fidelity of role-playing agents measured by psychological interviews (InCharacter, 32 characters, 14 scales). State-of-the-art agents matched human-perceived character personalities with up to 80.7% accuracy. What it can't see: Personality tests, not gameplay. | NPC Fidelity |
| RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models | Benchmarking and improving role-playing ability (RoleBench, 168,093 samples, 100 roles). Role-conditioned tuning let open models approach GPT-4 role-play. What it can't see: 2023 models; speaking-style focus. | NPC Fidelity |
Game state, rules and the AI game master
Whether an AI can run a game rather than tell a story — tracking state, enforcing rules, and assisting or replacing a game master.
| Paper | What it measures and found | Bears on |
|---|---|---|
| Orchestrated Reality: From Role-Play to Living, Playable Game Worlds -- LLM-Driven World Simulation as a Parameterized-Action POMDP | Architecture for LLM-run game worlds where the world is a canonical object owned by an orchestrator (work in progress). Argues that deployed systems let the narrator assert state in free prose "without any validated representation", so a fully autonomous engine remains infeasible. What it can't see: Work in progress; no evaluation yet. | Mechanical DepthMemory & ContinuityDeterminism & Fairness |
| IVIE: A Neuro-symbolic Approach to Incremental and Validated Generation of Interactive Fiction Worlds | Generating complete, playable interactive-fiction worlds with symbolic validation (IVIE). Symbolic validation grounded the LLM without removing creative freedom, though model inconsistencies occasionally bypassed puzzles. What it can't see: Generation of IF worlds, not live campaigns. | Mechanical DepthDeterminism & Fairness |
| World-State Transformations for Neuro-symbolic Interactive Storytelling | Neuro-symbolic interactive storytelling where the LLM triggers pre-programmed world-state changes. World-state transformations kept the world consistent while still encouraging creative player input. What it can't see: Eight participants; exploratory. | Mechanical DepthDeterminism & FairnessPlayer Agency |
| From World-Gen to Quest-Line: A Dependency-Driven Prompt Pipeline for Coherent RPG Generation | A staged, schema-enforced prompt pipeline for generating RPG worlds, NPCs and quests. Structured JSON hand-offs between stages reduced drift and kept content valid as complexity grew. What it can't see: Qualitative evaluation. | Mechanical DepthSignature Design |
| First Steps Towards Overhearing LLM Agents: A Case Study With Dungeons & Dragons Gameplay | "Overhearing" agents that listen to a D&D table and help the DM in the background. Some large audio-language models can perform background assistance from implicit audio cues. What it can't see: Workshop paper; assistive, not autonomous. | Mechanical Depth |
| STORY2GAME: Generating (Almost) Everything in an Interactive Fiction Game | Generating interactive-fiction games including the action code, with new actions created on demand (STORY2GAME). Generating action preconditions and effects lets stories stay open-ended yet grounded in tracked game state. What it can't see: Generated games, not live play. | Mechanical DepthPlayer Agency |
| PAYADOR: A Minimalist Approach to Grounding Language Models on Structured Data for Interactive Storytelling and Role-playing Games | Grounding an LLM on a minimal world representation by predicting action outcomes (PAYADOR). Predicting outcomes instead of mapping input to fixed actions preserves player freedom in RPG-style play. What it can't see: Minimal prototype. | Mechanical DepthDeterminism & FairnessPlayer Agency |
| RPGBENCH: Evaluating Large Language Models as Role-Playing Game Engines | LLMs as text RPG engines: creating a valid game (event-state representation) and simulating play while updating state and enforcing rules (RPGBench). State-of-the-art LLMs produce engaging stories but often fail to implement consistent, verifiable game mechanics, especially in long or complex scenarios. What it can't see: Models tested as engines, not shipped platforms with external state. | Mechanical DepthDeterminism & FairnessMemory & Continuity |
| You Have Thirteen Hours in Which to Solve the Labyrinth: Enhancing AI Game Masters with Function Calling | AI game masters using function calling for game-specific controls in a tabletop RPG (Labyrinth: The Adventure Game). Function calling improved both narrative quality and state-update consistency, by human evaluation and unit tests. What it can't see: One game system; workshop paper. | Mechanical DepthDeterminism & FairnessMemory & Continuity |
| LARP: Language-Agent Role Play for Open-World Games | Cognitive architecture for role-playing agents in open-world games (LARP). Combines memory processing, learnable action space and personality alignment for open-world NPCs. What it can't see: Framework paper. | Memory & ContinuityNPC Fidelity |
| CALYPSO: LLMs as Dungeon Masters' Assistants | LLM tools that assist a human Dungeon Master (CALYPSO). DMs found the output good enough to read to players and used it for ideas while keeping creative control. What it can't see: Assists a human DM rather than replacing one. | Mechanical DepthSignature Design |
| FIREBALL: A Dataset of Dungeons and Dragons Actual-Play with Structured Game State Information | D&D actual play with true game state from ~25,000 Discord sessions using the Avrae bot (FIREBALL). Giving models real game-state information improved generated turns; models can also produce executable game commands. What it can't see: Human play logs; model generation evaluated, not full GM play. | Mechanical DepthDeterminism & Fairness |
| Rolling the Dice: Imagining Generative AI as a Dungeons & Dragons Storytelling Companion | Design vision for generative AI as a D&D storytelling companion. Proposes design guidelines and flags immersion and cognitive-load questions. What it can't see: Position paper, no experiment. | Player AgencySignature Design |
| Dungeons and Dragons as a Dialog Challenge for Artificial Intelligence | D&D as a dialogue challenge: generating the next turn and predicting game state (~900 games, 800,000 turns). Tracking game state measurably improves the model's turns; frames D&D as a benchmark problem for AI. What it can't see: Pre-ChatGPT models. | Mechanical DepthNPC Fidelity |
Player agency and narrative control
How much a player can change the story, and how systems keep authored structure without taking that away.
| Paper | What it measures and found | Bears on |
|---|---|---|
| Design Techniques for LLM-Powered Interactive Storytelling: A Case Study of the Dramamancer System | Design techniques for turning author-written story schemas into player-driven play (Dramamancer). Frames the author-intent vs player-agency balance as a design problem. What it can't see: Extended abstract. | Player AgencySignature Design |
| SNAP: A Plan-Driven Framework for Controllable Interactive Narrative Generation | Plan-driven control of interactive narratives to prevent drift (SNAP). Splitting the story into planned cells kept dialogue scenario-consistent under varied user input. What it can't see: Short paper. | Player AgencyMemory & Continuity |
| WHAT-IF: Exploring Branching Narratives by Meta-Prompting Large Language Models | Branching interactive fiction generated from a linear story by meta-prompting (WHAT-IF). Storing the branch tree as a graph kept alternate storylines coherent. What it can't see: Choice-based IF, not free text. | Player Agency |
| Player-Driven Emergence in LLM-Driven Game Narrative | Emergent narrative from player interaction with GPT-4 NPCs in a mystery text adventure (28 players). Players discovered new story nodes not in the original narrative; the most "emergent" players were those who like exploration games. What it can't see: One game, small sample. | Player AgencySignature Design |
| GENEVA: GENErating and Visualizing branching narratives using LLMs | LLM tool that generates branching and reconverging narrative graphs for designers (GENEVA). GPT-4 produced rich branching narratives under designer constraints. What it can't see: Authoring tool, not a runtime GM. | Player AgencySignature Design |
AI that plays games (context)
Benchmarks where the AI is the player, not the game master. Useful context, but they measure a different problem from ours.
| Paper | What it measures and found | Bears on |
|---|---|---|
| ByteSized32Refactored: Towards an Extensible Interactive Text Games Corpus for LLM World Modeling and Evaluation | Corpus of 32 text games for LLM world modelling (ByteSized32Refactored). Refactored, extensible text-game corpus for generation and evaluation. What it can't see: Resource paper. | — |
| Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games | LLM agents playing 12 popular video games across genres (Orak). Plug-and-play benchmark and fine-tuning data for game-playing agents. What it can't see: Measures AI as the player, not as the game master. | — |
| TALES: Text Adventure Learning Environment Suite | LLMs playing synthetic and human-written text adventures (TALES). Even top agents score under 15% on text adventures designed for human enjoyment. What it can't see: AI as player. | Memory & Continuity |
| TextArena | Competitive text games for agent evaluation (TextArena, 57+ environments). Online leaderboard with TrueSkill for social skills like negotiation and deception. What it can't see: AI as player. | — |
| DSGBench: A Diverse Strategic Game Benchmark for Evaluating LLM-based Agents in Complex Decision-Making Environments | Strategic decision-making in six complex games (DSGBench). Fine-grained scoring of agents in long-horizon strategy games. What it can't see: AI as player. | — |
| Can VLMs Play Action Role-Playing Games? Take Black Myth Wukong as a Study Case | Vision-language agents playing Black Myth: Wukong. Explores playing an action RPG from screen input only. What it can't see: Combat play, not narrative. | — |
| Towards a Holodeck-style Simulation Game | A Holodeck-style simulation game built on generative agents (Infinitia). Applies the Generative Agents idea to a playable, multiplayer Unity sandbox. What it can't see: System description. | Signature Design |
| Generative Agents: Interactive Simulacra of Human Behavior | Believable agents with memory, reflection and planning in a Sims-like town (Generative Agents). The foundational architecture — store experiences, reflect, retrieve to plan — behind most "AI NPC" work since. What it can't see: Social simulation, not RPG play. | NPC FidelityMemory & Continuity |
| Can Large Language Models Play Text Games Well? Current State-of-the-Art and Open Questions | How well ChatGPT plays text games (2023). ChatGPT could not build a world model from play or even the manual. What it can't see: Early model generation. | Memory & Continuity |
Surveys and field maps
Reviews that map the role-playing agent literature as a whole.
| Paper | What it measures and found | Bears on |
|---|---|---|
| Role-Playing Agents Driven by Large Language Models: Current Status, Challenges, and Future Trends | 2026 review of role-playing agents: personality modelling, memory, evaluation. Traces the field from templates to cognitive simulation. What it can't see: Review. | NPC FidelityMemory & Continuity |
| Towards a Design Guideline for RPA Evaluation: A Survey of Large Language Model-Based Role-Playing Agents | Evaluation-design guideline from a review of 1,676 role-playing agent papers. Identifies six agent attributes, seven task attributes and seven evaluation metrics used across the literature. What it can't see: Guideline for researchers. | NPC Fidelity |
| The Oscars of AI Theater: A Survey on Role-Playing with Language Models | Survey of role-playing with language models ("The Oscars of AI Theater"). Taxonomy of data, models, architecture and evaluation for role-play. What it can't see: Review. | NPC Fidelity |
| From Persona to Personalization: A Survey on Role-Playing Language Agents | Survey of role-playing language agents: demographic, character and individualised personas. The most-cited field map of role-playing agents. What it can't see: General role-play, games are one application. | NPC FidelityMemory & Continuity |
| A Survey on the Memory Mechanism of Large Language Model based Agents | Survey of memory mechanisms in LLM-based agents: design and evaluation of memory modules. The standard reference for how agent memory is built and evaluated. What it can't see: General agents, not games. | Memory & Continuity |
In Industry: Where These Ideas Ship
The research above shows up in shipped products — middleware for game studios, mods, and games on Steam that run language models live. These entries are descriptive and unrated: we have not played them for the benchmark, so we say what each one is, in its maker's own words, and nothing about how good it is. For a player's guide to the games, see games with AI NPCs you can talk to; for NVIDIA's character toolkit, NVIDIA ACE explained; for the two best-known character platforms, Convai and what happened to Inworld AI.
| Name | Status | What it is |
|---|---|---|
| NVIDIA ACE | shipping (SDK) | NVIDIA describes ACE as "a suite of digital human technologies that power agentic workflows for autonomous game characters and digital assistants." |
| Inworld AI | pivoted | Once known for game-character AI; its homepage now sells "Realtime TTS and STT models, LLM serving, and the inference behind both" for consumer-facing applications. |
| Convai | shipping | Positions itself as "Conversational AI for Virtual Worlds." |
| Ubisoft NEO NPC (R&D) | research prototype | A small R&D team at Ubisoft Paris experimenting with generative AI toward real conversations with NPCs, per Ubisoft's own news post. |
| Mantella | active (repo pushed 2026-07-16) | A Skyrim and Fallout 4 mod "which allows you to naturally speak to NPCs using a Speech-to-Text → LLMs → Text-to-Speech pipeline." |
| AI Roguelite | released 2023-10-25 | Developer disclosure on Steam: it "heavily uses AI to live-generate in-game content such as text, images, and sound effects" and "to make a variety of game mechanics decisions in real time." |
| inZOI | early access since 2025-03-27 | Steam disclosure: player text can influence "character actions and thoughts … using SLM technology", plus AI-generated textures, 3D objects and motions. |
| Where Winds Meet | released 2025-11-14 | Steam disclosure: "AI-driven NPC text and voice chat that responds to player input in real time." |
| Suck Up! | released 2025-10-01 | Steam disclosure: players talk to AI characters by voice and "the AI responds in real time based on tone and strategy." |
| Wanderfolk | coming 2026 | Steam disclosure: villager conversations "generated in real time by a large language model", and NPCs "remember what you've said to them across the playthrough." |
| Vaudeville | released 2025-11-28 | Steam disclosure: "dialogues generated with the help of a conversational AI"; players talk to townsfolk by typing or voice. |
How This Index Is Built
- What gets in. Research that bears on at least one of our seven axes, or on how an AI RPG is built. We searched arXiv for role-playing, game-master, text-adventure, interactive-fiction and NPC work, read every candidate's abstract, and kept 63 of about 255. Papers where the AI is the player are kept as context, in their own group.
- How it is checked. Titles, authors and dates come from the arXiv record, not from memory or from another summary. A script re-checks every title against arXiv and every industry link against its source before each update.
- Venues. We show a conference or journal only when the authors state it on arXiv. A paper that says it was submitted somewhere is listed as a preprint, because submitted is not accepted.
- Freshness. The index is re-reviewed every quarter, alongside each benchmark deposit. New papers that bear on the axes are added; nothing is quietly removed.
- Corrections. If we have misread a paper, tell us. Mistakes we fix are logged on the corrections page, like any other.
Take the Data
The whole index is available as research.json under CC BY 4.0: every entry, its source, its axis mapping and the date it was verified. Our summaries and mappings are free to reuse with attribution. The papers belong to their authors — we link to them rather than reproduce them.
For the measurements this index is set against, see the benchmark. For how those measurements are made, see the methodology.