How AI Companion Content Filters Actually Work

People treat "the filter" as a single thing with a single setting. It is closer to four or five independent systems stacked on top of each other, each with its own rules, owner and failure mode. Understanding the layers explains most of the behaviour that otherwise looks arbitrary: why one platform allows something another blocks, why the same app is inconsistent between features, and why a filter sometimes fires on a completely innocent message.

This page describes mechanisms only. It does not test or recommend ways around any platform's rules.

Filters are layered, not singular

A message travelling through a companion app can be checked at several points, and each check is a separate decision.

Layer Where it sits What it typically decides
Model alignment Inside the language model itself Whether the model will engage with the topic at all, in the register it was trained to use
Input classification Between you and the model Whether your message is allowed to reach the model
Output classification Between the model and you Whether the generated reply is allowed to be shown
Application policy Product rules, often per-feature Age gates, specific scenarios, which features are stricter than others
Commercial and platform rules App stores, payment processors, hosting Constraints the product accepts to keep its distribution and payment rails

The important point: only one of these lives inside the model. When a platform is described as having "a strict model", it usually means the platform configured the outer layers strictly, not that it is running a different model.

Why the same model behaves differently on two platforms

Four reasons, and they are not equivalent.

That last point is the one users trip over most: the filter is not one setting, so "the app blocked me" is an incomplete description. Which surface, and which layer, matters.

The two failure modes

Filters are classifiers, and classifiers are wrong in both directions.

False positives. Benign content is blocked because it statistically resembles something disallowed. Common causes: context the classifier cannot see (a scene that reads very differently without the previous twenty messages), medical or clinical vocabulary, discussion of the topic rather than depiction of it, and non-English text being classified by a model trained mostly on English. False positives are the main reason users report inconsistent behaviour — the same message can pass in one context and fail in another.

False negatives. Content that should be caught is not. This is normal and expected: classifiers are probabilistic, and there is no clean boundary between allowed and disallowed in natural language. It also means a platform's advertised filter is a policy statement about intent, not a guarantee about outcome.

Neither failure is a sign of incompetence. The engineering problem is genuinely hard, and the trade-off between them is a product decision: strictness reduces one kind of error and increases the other.

Why filter rules change over time

Rules move in both directions, and usually for reasons unrelated to the users.

Because the layers belong to different parties, a change you notice may have been made by someone who is not the platform at all — a provider updating its classifier, or a processor changing its terms.

Age gates are the weakest layer

Most companion apps state a minimum age in their terms, and most enforce it with a self-declared date of birth, sometimes with an additional confirmation prompt. That is a compliance mechanism, not a technical barrier.

For parents, the checks that actually carry weight are: the store rating and its stated age band, what payment method is attached to the account, and whether the product offers any parental controls at all. A filter being advertised is not evidence that it is enforced for a given account.

Why some features are filtered harder than others

Look at a companion app's feature list and you will usually find the same gradient: text chat is the most permissive, images and voice are stricter, public and shared surfaces are strictest.

There are two reasons. Media is harder to classify reliably, so the cheapest safe policy is to restrict it. And public surfaces are the ones that create the platform's legal and reputational exposure — a private conversation is a user's problem, a public gallery is the platform's.

This is why a subscription can look generous on messaging and restrictive on image and voice generation: the quotas and the filters are usually decided together.

What you can legitimately check before paying

  1. Read the policy for scope, not for adjectives. Three questions: does it distinguish private from public content, does it distinguish text from media, and does it say who decides? A policy that answers all three is usually written by someone who understands the product.
  2. Check the pricing page for what is quota-limited. Per-generation limits are frequently where the practical restrictions live.
  3. Look at the terms for the payment and cancellation path before you subscribe, not after. Our subscription traps page covers the specific patterns worth checking, and the terms checklist tool walks through them as a printable list.
  4. Test on the cheapest tier. Filter behaviour, like memory behaviour, is often what the paid tier actually changes.

Where this leaves you

Filters are a normal part of a product built on top of a hosted model, and they are not going away. The realistic expectations are: rules will differ between platforms and between features on the same platform, they will change without notice, and they will occasionally be wrong in both directions.

The practical approach is to treat the filter as part of the product's design — something to evaluate when choosing where to spend money, rather than an obstacle to be defeated. Products whose filter policy you can read and predict are easier to live with than ones that surprise you, even when the readable one is stricter.

Key takeaways

For the vocabulary, see the AI companion glossary. For why a scene can collapse mid-conversation without any filter being involved, see why characters break character.