LIVECREATION OS® WORLD-MANUFACTURING SYSTEMEST. 2026 · PERTH, WESTERN AUSTRALIAPOWERED BY CANONLOCK® IISYS v4.0

Test · The memory probe

A memory test you can run on 6 AI story games.

Search “how long can an ai campaign last” and you get broad claims that are hard to compare. So I wrote down a test simple enough that you can run it yourself, on whatever you already play, in about twenty minutes. The service descriptions below are not controlled head-to-head results.

Why a memory test, and why this one

An AI story game only feels alive if the world holds a grudge, keeps an inventory, and enforces a debt. Those three things are the difference between a campaign and a very good improv partner who resets every scene. The trouble is that “it remembers” is unfalsifiable marketing until you define what remembering means and put a number on how long it lasts.

I want to be honest about the scope of this piece up front, because I did not run controlled multi-thousand-turn sessions inside every competitor's account. Several are invite-gated or writing-forward rather than campaign-forward, and pretending I have clean lab data for each would be dishonest. What I can do is two things that are true: define a probe precise enough that anyone can run it, and describe how each tool is designed to hold state and what each maker publicly says it stores. Creation OS separately publishes one world artifact sealed at turn 5,011. That record is not a controlled comparison with the other five services.

The probe: grudge, item, debt

The test has three parts. Each is a small, concrete fact that a living world should carry forward without you re-stating it. Run them early, keep playing dense turns, then come back and check.

  • The grudge. Wrong a specific, named character early. Rob them, insult them, break a promise. Note the name. Play at least thirty more turns of unrelated events, then walk back into that character's space. Do they treat you as the person who wronged them, by name, without you reminding them?
  • The item. Acquire or sell one specific object with a distinct name. A brass compass. A grey mare called Ash. Twenty-plus dense turns later, ask where it is or try to use it. Does the world agree with the ledger, or does it improvise a fresh answer?
  • The debt. Owe an exact amount to a named party. Forty coins to the harbourmaster. Keep playing. Later, try to walk past that party as if the debt never happened. Is the number still exact, and is it actually enforced against you, or does it quietly evaporate?

The reason those three probes are useful is that they attack the seam between fluent callback and durable state. A world that holds all three after a long interval has produced useful evidence, but one pass does not establish unlimited memory. The probe is deliberately easy to repeat across different kinds of state.

The six tools at a glance

Here is the honest version of the comparison table. The middle column is how each tool is built to carry state, in the makers' own framing. The right column tells you what the probe still needs to establish. It is not a lab verdict for any row.

ToolHow it holds stateWhat to verify with the probe
AI DungeonFinite generation context plus automatic summaries, a Memory Bank, embeddings, Story Cards, and Plot Components.Whether an incidental fact can be retrieved without first turning it into an authored card or component.
NovelAIA writing-forward tool with a Lorebook you author, injected into context when its keys are triggered.Whether an exact number or one-off object remains correct without a dedicated Lorebook entry.
Friends & FablesStructured campaign objects, separate narration and state contexts, lore, and relevance-gated long-term memory.How reliably each kind of fact returns at the interval and campaign length that matter to you.
VoyageA deterministic world and state system, closer in philosophy to a standing record than to raw context.Whether its current product access lets you repeat the probe and inspect the resulting state yourself.
ChatGPT / plain LLMA fixed control: a new plain chat with product memory and external tools disabled, using only the supplied transcript.How the result changes as the transcript grows. This control does not represent ChatGPT with its full product memory enabled.
Creation OSStructured records for supported world and mechanical state, plus episodic and semantic retrieval for narrative context.Whether the tested fact maps to supported state, how the Narrator handles it, and whether any correction is required.

AI Dungeon: automatic and authored memory

AI Dungeon is the tool most people mean when they say “ai rpg,” and it is genuinely flexible. Its current continuity system includes automatic summaries, a Memory Bank, embeddings, Story Cards, and Plot Components alongside the generation context. It is not accurate to describe the service as a bare chat transcript.

The useful question is how the automatic systems handle facts you did not deliberately author. Run the grudge, item, and debt probe with the current default settings, then repeat it with Story Cards or Plot Components. That shows both the automated result and the amount of curation you prefer.

NovelAI: writer-controlled context

NovelAI is excellent at what it is for, which is writing. Its Lorebook lets you author entries that get injected into context when their trigger words appear, so a well-maintained Lorebook can carry a named character or a recurring place a long way. If you treat it like a bookkeeping system and feed it, it rewards you.

It is prose-first by design, so the probe should distinguish raw transcript performance from a carefully maintained Lorebook. An exact debt and an incidental one-off item are useful tests because they are easy to verify and may not receive an authored entry by default.

Friends & Fables: structured campaign memory

Friends & Fables is interesting here because memory is explicitly part of the product. It uses structured campaign objects, separate narration and state contexts, lore, and relevance-gated long-term memory. That is a serious state and retrieval architecture, not a prompt attached to a chat window.

We did not run a controlled, long-form account test that supports a specific failure point. Run the same probe at the interval and campaign length that matter to you. Record whether the result came from structured state, retrieved memory, an authored lore entry, or a correction you supplied.

Voyage: the right philosophy, behind a gate

Voyage is the one I want to be most careful about, because it is closest to the approach I believe in. It is described as a deterministic world and state system, which means it is trying to hold facts as facts rather than lean on the context window. On paper, that is exactly the design that passes a grudge, item, and debt probe.

We did not collect a controlled endurance result for Voyage, so this page assigns it no pass, fail, or turn count. If you have access, run the same probe and note whether you can inspect the stored state behind the generated prose.

ChatGPT and plain LLMs: the honest baseline

A fixed plain-chat setup is the control group: start a new chat, turn off product memory and external tools, and supply only the running transcript. This deliberately does not represent the full ChatGPT product. For a first session this feels magical, because everything you have said recently is right there and the model weaves it beautifully.

Do not assume a failure at turn 20, 50, or any other fixed number. Context size, turn density, prompt length, and model choice all change the result. The value of the control is that its memory source is explicit, so another person can reproduce the setup.

The Creation OS record you can inspect

Here is why Creation OS is the strong row, and I want to be precise about the reason. It is not that the Narrator is smarter. Every tool here uses capable models. The distinction is that supported world and mechanical state can be held in structured records apart from the prose, while narrative memory uses layered retrieval. A debt or item must map to a supported system before it can be mechanically enforced.

Creation OS publishes one world artifact sealed at turn 5,011, with an event record available at creationos.io/canonlock. You do not have to trust me. You can inspect that artifact and the method used to verify it. It is evidence from one long run, not a controlled result against the other five services and not a promise that every narrative fact survives.

I will keep my own honesty rule here too. The Narrator can still slip on small wording in the prose, the same way any model can, and when it does, an out-of-character correction can repair the record. Structured mechanics remain a stronger source of truth than generated wording, but only for systems the product actually tracks.

Run the probe yourself

The honest takeaway is the one I would want if I were you: do not trust any roundup, including this one, more than a test you ran with your own hands. So here is the whole method in one paragraph, ready to use on whatever you play tonight.

  • Pick a named character in the first few turns and wrong them on purpose. Write the name on a sticky note.
  • Acquire or sell one specific, distinctly named object. Write that down too.
  • Take on an exact debt to a named party. Write the number.
  • Play thirty to fifty dense, unrelated turns. Do not remind the world of any of the three.
  • Come back. Face the character, ask about the object, try to skip the debt. Score each: held by name, held vaguely, or gone.

Run that on AI Dungeon, on NovelAI, on Friends & Fables, on a plain chat control, or anything else you like. Then run it on Creation OS and compare your result with its published 5,011-turn artifact. Whatever wins your own repeated tests is the better fit for your campaign.

RUN THE SAME PROBE

WORLD RECORD
CONSEQUENCES / ON FILE
SCENE HOLD®
CATCH-UP / ON RETURN
MECHANICS
MONEY · GEAR · STANDING
GENRE RANGE
FOOTBALL · NOIR · COZY · FANTASY
Run the test yourself

Free tier. First world on the house.