How AI Companion Memory Actually Works
Memory is the feature people judge companion apps on, and the one they understand least. The short version: a language model does not remember anything. It has no storage, no diary and no thread of continuity between sessions. What feels like memory is text that a piece of software decided to put in front of the model again. This page explains the mechanisms, so you can tell a good memory system from a demo.
There is no memory, only a window
A model receives a block of text and produces the next piece of text. That is the whole operation. When people talk about a model "remembering", what actually happened is that a previous message was included in the block again.
So the useful question is never "does it remember?" but "what does it decide to include, and what does it throw away?" Everything below is a variation on that decision.
The context window is a budget, not a container
The block of text is limited by the model's context window, measured in tokens. Tokens are not words; a rough working rule is that a token is shorter than a word, so a long chat is more expensive than it looks.
The window has to fit all of this at once:
- the character definition and any system instructions
- whatever the app decided to include as "memory"
- your recent messages
- the character's own recent replies
- room for the reply currently being generated
These compete. A long persona file leaves less room for history. A long history leaves less room for the persona. When the window fills up, something has to go — and what goes first is a product decision, not a technical necessity.
Character cards are the clearest example of the trade: a card is a fixed cost paid on every single turn, forever.
Four ways apps handle overflow
Almost every companion product is one of these, or a combination.
| Design | What happens when the window fills | Cost | Weakness |
|---|---|---|---|
| Rolling window | Oldest messages are dropped | Cheapest | Anything not recently said is gone |
| Summary compaction | Old messages are replaced by a summary written by a model | Moderate | The summary is a lossy rewrite; details and nuance vanish |
| Fact extraction plus retrieval | Specific facts are stored separately and pulled back in when they seem relevant | Higher | Facts can be missed, mis-scoped, or retrieved for the wrong reason |
| Very large window, no compaction | Everything is kept | Most expensive per message | Still finite; quality can drop in the middle of very long inputs |
Most platforms combine the middle two: summarise most of the history, keep a small store of durable facts, and always keep the last handful of messages verbatim. Which combination an app uses is the single biggest reason two products using the same class of model feel different.
Why the character forgets something you told it
Not one cause — five, and they look identical from the outside.
- It aged out. The message was dropped from the rolling window. This is the most common one, and it is not a bug.
- The summary lost it. A model rewrote forty messages into three sentences. Your detail was not in those sentences.
- It was never extracted. Fact extraction is selective. "My sister is called Ana" might be stored; "I have been sleeping badly since my sister Ana moved away" may yield one fact and lose the rest.
- Retrieval did not fire. Fact stores usually retrieve by similarity to what you just said. If you refer to it indirectly, the matching step can miss.
- It was there and got ignored. Long inputs are not attended to uniformly; material buried in the middle of a very long context is used less reliably than material at the start or end. The information was present in the window and still did not affect the reply.
Only the last one is a model-level effect. The first four are engineering decisions you can partly infer from behaviour.
Why it invents memories
Models generate plausible continuations. When a detail is missing, the cheapest way to produce fluent text is to produce a plausible one. This is confabulation, not deception, and it gets worse in exactly the situations people care about most: emotional, specific, personal questions.
Summarisation compounds it. A summary is itself generated text, so a slightly wrong summary written once can be repeated confidently for weeks.
The practical consequence: anything important should be verifiable, not remembered. Keeping your own notes outside the app is not paranoia, it is the only copy that survives.
What memory costs the platform
"More memory" is a trade, not a straight upgrade, which is why it is usually a paid tier.
- Money. Every retained token is paid for on every turn until it is dropped. A memory that doubles the prompt roughly doubles the per-message cost.
- Latency. A longer prompt takes longer before the first token appears. A retrieval step adds a round trip before generation even starts.
- Quality. Past a point, more history can make replies worse. Irrelevant context dilutes the relevant context. This is why some apps deliberately keep memory small and curated.
The best implementations are not the ones that store the most. They are the ones that store the right things and can show you the list.
How to test a memory system before you pay
Do not trust a feature name. Run this instead:
- Plant a specific fact — a name, a street, a food you dislike — and ask about it after twenty or thirty messages, then again the next day. Same detail, different session, is the real test.
- Contradict yourself. Tell it one thing, then the opposite a day later, and see which it keeps. A product with a fact store often keeps both; a summarising product usually keeps the newer one.
- Ask it what it remembers about you. The answer is usually close to the actual memory block. That single question tells you more than the marketing page.
- Check whether memory is user-editable. Being able to read and delete what was stored is the difference between a database and a black box. It also tells you what the privacy policy is dealing with.
- Look for the memory limit in pricing. "Unlimited messages" with a small memory is a different product from "unlimited memory".
If you are building a character yourself, the same mechanisms are what your character card is fighting against: a card is fixed text that competes with history for the same budget.
Key takeaways
- A model has no memory. Anything the character "remembers" was re-inserted into the context window by the app.
- The window is a shared budget between the character definition, stored memory and recent messages. Something always loses.
- Rolling windows, summary compaction and fact stores are the three standard designs; most products combine them.
- Forgetting has five distinct causes, and only one of them is a model limitation.
- Models confabulate missing details fluently, and summaries can bake in an error permanently.
- More memory costs money and latency per message, and can reduce quality if it dilutes the relevant context.
- Test memory by planting a specific fact and checking it across sessions, not by reading feature lists.
For the vocabulary used here, see the AI companion glossary. For the wider product picture, start with what an AI companion is.