How AI Companion Memory Actually Works

Memory is the feature people judge companion apps on, and the one they understand least. The short version: a language model does not remember anything. It has no storage, no diary and no thread of continuity between sessions. What feels like memory is text that a piece of software decided to put in front of the model again. This page explains the mechanisms, so you can tell a good memory system from a demo.

There is no memory, only a window

A model receives a block of text and produces the next piece of text. That is the whole operation. When people talk about a model "remembering", what actually happened is that a previous message was included in the block again.

So the useful question is never "does it remember?" but "what does it decide to include, and what does it throw away?" Everything below is a variation on that decision.

The context window is a budget, not a container

The block of text is limited by the model's context window, measured in tokens. Tokens are not words; a rough working rule is that a token is shorter than a word, so a long chat is more expensive than it looks.

The window has to fit all of this at once:

These compete. A long persona file leaves less room for history. A long history leaves less room for the persona. When the window fills up, something has to go — and what goes first is a product decision, not a technical necessity.

Character cards are the clearest example of the trade: a card is a fixed cost paid on every single turn, forever.

Four ways apps handle overflow

Almost every companion product is one of these, or a combination.

Design What happens when the window fills Cost Weakness
Rolling window Oldest messages are dropped Cheapest Anything not recently said is gone
Summary compaction Old messages are replaced by a summary written by a model Moderate The summary is a lossy rewrite; details and nuance vanish
Fact extraction plus retrieval Specific facts are stored separately and pulled back in when they seem relevant Higher Facts can be missed, mis-scoped, or retrieved for the wrong reason
Very large window, no compaction Everything is kept Most expensive per message Still finite; quality can drop in the middle of very long inputs

Most platforms combine the middle two: summarise most of the history, keep a small store of durable facts, and always keep the last handful of messages verbatim. Which combination an app uses is the single biggest reason two products using the same class of model feel different.

Why the character forgets something you told it

Not one cause — five, and they look identical from the outside.

Only the last one is a model-level effect. The first four are engineering decisions you can partly infer from behaviour.

Why it invents memories

Models generate plausible continuations. When a detail is missing, the cheapest way to produce fluent text is to produce a plausible one. This is confabulation, not deception, and it gets worse in exactly the situations people care about most: emotional, specific, personal questions.

Summarisation compounds it. A summary is itself generated text, so a slightly wrong summary written once can be repeated confidently for weeks.

The practical consequence: anything important should be verifiable, not remembered. Keeping your own notes outside the app is not paranoia, it is the only copy that survives.

What memory costs the platform

"More memory" is a trade, not a straight upgrade, which is why it is usually a paid tier.

The best implementations are not the ones that store the most. They are the ones that store the right things and can show you the list.

How to test a memory system before you pay

Do not trust a feature name. Run this instead:

  1. Plant a specific fact — a name, a street, a food you dislike — and ask about it after twenty or thirty messages, then again the next day. Same detail, different session, is the real test.
  2. Contradict yourself. Tell it one thing, then the opposite a day later, and see which it keeps. A product with a fact store often keeps both; a summarising product usually keeps the newer one.
  3. Ask it what it remembers about you. The answer is usually close to the actual memory block. That single question tells you more than the marketing page.
  4. Check whether memory is user-editable. Being able to read and delete what was stored is the difference between a database and a black box. It also tells you what the privacy policy is dealing with.
  5. Look for the memory limit in pricing. "Unlimited messages" with a small memory is a different product from "unlimited memory".

If you are building a character yourself, the same mechanisms are what your character card is fighting against: a card is fixed text that competes with history for the same budget.

Key takeaways

For the vocabulary used here, see the AI companion glossary. For the wider product picture, start with what an AI companion is.