Why AI Characters Break Character in Long Conversations
Everyone who talks to a companion app long enough hits the same wall. The character is sharp for fifty messages, then says something that is recognisably not them. This page covers what is actually happening, which causes are common, and which fixes are worth the effort.
What "breaking character" actually looks like
It is not one failure. Five distinct problems get called "breaking character", and they have different causes.
- Drift. The voice flattens over time. The character keeps its name and loses its mannerisms.
- Contamination. Assistant voice bleeds through: helpful, neutral, apologetic, prone to lists.
- Contradiction. Facts or traits established earlier get overwritten without comment.
- Repetition. The same sentence pattern, gesture or greeting returns every few turns.
- Refusal. The character stops being a character and becomes a policy statement, usually in the app's voice rather than its own.
If you know which one you are seeing, most of the diagnosis is already done.
Cause 1: the persona loses the budget fight
The character definition is fixed text. The conversation is not. Every turn, the growing history pushes against the fixed persona for room in the context window — and something has to be cut.
The failure mode is quiet: the app summarises the history, the persona gets truncated, or the system instruction is placed where it is least attended to. The character does not announce that it lost half its personality file. It just gets blander.
This is the most common cause, and the reason a character often feels best in the first session and worst after a long gap. See how companion memory works for the mechanics of what gets dropped.
Cause 2: the instructions contradict each other
Character files accumulate rules. Written over weeks, they end up like this: reserved, does not open up easily; warm and direct; uses long poetic sentences; keeps replies under two sentences.
A model cannot satisfy a contradiction. What it does instead is average the conflicting instructions, which produces a character that is vaguely pleasant and has no edges. Contradictions do not cancel out; they dilute.
This is the cause you have the most control over, and the one most people never check.
Cause 3: the assistant prior never fully goes away
Every chat model was trained to be a helpful assistant: answer the question, be accurate, avoid harm, decline politely when asked. The character definition is a thin layer on top of that.
Under pressure, the underlying training shows through. Pressure includes: an unusual request, a long ambiguous scene, a topic the model's safety training treats as sensitive, or a prompt that accidentally reads like a task ("summarise what I said"). The reply that comes back is an assistant's reply with the character's name attached.
This is not the model failing to understand. It is the model reverting to what it was most heavily trained to do. Long character cards with example dialogue help because they give the model a concrete pattern to imitate instead of a description to follow.
Cause 4: summarisation rewrites the character
When an app compacts history, it usually uses a model to write the summary. That summary is generative text, and it inherits the same bias as the chat itself.
Three things get lost in the rewrite, in this order:
- Register. Speech style is the first casualty; summaries record facts and discard tone.
- Relationship state. How close the characters are, what was settled, what is still unresolved.
- The uncomfortable parts. Anything that made the character difficult — stubbornness, an argument — tends to get smoothed over, because a summary aims for coherence.
The character then continues from a version of itself that has been edited into someone more agreeable. It is not the model drifting; it is the record being rewritten.
Cause 5: long context is not uniformly attended to
Material in the middle of a very long input is used less reliably than material at the start or the end. This is a known property of how these models process long inputs, not a bug in any one product.
The practical effect: instructions placed in the middle of a large context — including persona details inside a long card, or a rule stated forty messages ago — can be present in the window and still not influence the reply. When the same character is consistent in short sessions and inconsistent in long ones, this is often why.
Cause 6: the safety layer interrupts
Some breaks are not the character at all. They are a filter, an app-level policy rule, or the provider's moderation layer stepping in mid-scene. The tell is the register: policy text has a recognisable flat, formal tone, and it does not use the character's voice.
That is a different system from the one generating the persona, and no amount of card writing fixes it. Our filter mechanics page covers where those layers sit.
What actually reduces it
Ranked by how much difference they make:
- Remove contradictions from the character file. One clear set of traits beats four overlapping ones. This is free and it is the biggest single win.
- Add example dialogue. Concrete samples of the character speaking do more for consistency than any adjective. The guide to writing a character prompt covers how much to include.
- Keep scenes short. Long scenes are where context pressure and summarisation do the most damage. Starting a fresh scene preserves the character better than rescuing a collapsed one.
- Pin the non-negotiables. Keep the list of traits that must never drift short — two or three — and put them where they survive compaction.
- Use a stable, well-structured card. Structured fields survive rewriting better than prose. The character card builder exports the standard structure; cards explained covers why the fields are shaped that way.
- Re-state the relationship, not the personality. When a long chat starts to slip, reminding the character who you are to each other recovers more than re-describing its personality.
What does not help
- "The model is getting tired." There is no fatigue, no wear, no learning from your conversation. State does not persist between sessions beyond what the app chose to store.
- "It is learning about me." Nothing is being trained on your chat. Behaviour changes because the input changed, not because the model changed.
- Repeating the instruction louder. Emphasis on a contradictory instruction makes the averaging worse, not better.
- Threatening or pleading in the prompt. It changes tone, not capability. The character does not comply out of pressure.
The realistic ceiling
Perfect consistency across thousands of messages is not a solved problem. Any claim that a platform has eliminated drift should be treated as a marketing claim until you have tested it across sessions, with a character whose voice is specific enough to notice when it slips.
What good products do is degrade gracefully: keep the voice recognisable after the details have gone, and give you a way to reset the scene without losing the character. That is a reasonable standard to hold them to. Writing a character that survives it is a separate skill, covered in how to create an AI character.
Key takeaways
- Five different failures get called "breaking character"; they have different causes and different fixes.
- The most common cause is the persona losing its share of the context window to growing history.
- Contradictory traits do not cancel, they average — producing a bland character.
- Every chat model has an assistant prior underneath the persona, and pressure makes it show.
- Summary compaction rewrites the character in the direction of being agreeable, and tone is the first thing lost.
- Information in the middle of a long context is used less reliably than information at the edges.
- Removing contradictions and adding example dialogue are the two highest-leverage fixes; neither costs money.