Voice, Image and Video in AI Companions: Three Technical Routes

A companion app is usually described as one product. Technically it is a chat system with up to three unrelated generation pipelines attached: speech, images and video. Each is built on different technology, fails in different ways, and costs the platform different amounts of money. That is why the quotas and the quality vary so much between them.

One product, four systems

The text chat is a language model producing tokens. The other three are entirely separate:

They do not share code paths, they are not the same model, and on most platforms they do not share quota. When a subscription "includes voice, images and video", the useful question is which of the three is actually usable within the monthly allowance.

Route 1: speech

Two things get called voice, and they are different products.

Text-to-speech converts the character's reply into audio. Modern neural TTS is good enough to be mistaken for a recording in short sentences. The engineering problems are latency and streaming: audio has to start before the whole reply is generated, or the character feels slow to answer.

Voice cloning takes a sample of a voice and makes the model speak in it. The technical bar has fallen far enough that a short sample is often sufficient, which is why consent and misuse are the live issues here rather than quality.

Practical characteristics:

If voice quality matters to you, test it on the specific thing you will use it for: reading a long emotional reply, not a sample sentence.

Route 2: images

Image generation in this category is almost entirely diffusion-based: the model starts from noise and iteratively removes it, guided by your text.

The interesting problem is not generating a picture. It is generating the same person twice.

A text prompt describes a character in general terms, and a diffusion model has no memory of the last picture it made. Ask for the same character in a different scene and you get a different face. Three approaches are used to fight this:

Approach How it works Trade-off
Reference or style conditioning The previous image, or a reference image, is fed in alongside the prompt Simple; identity still drifts over repeated generations
Trained character adapters A small model is trained on a handful of images of one character, then loaded at generation time Much stronger consistency; needs a training step and per-character storage
Structured prompt discipline A fixed appearance block is reused verbatim in every prompt No extra cost; only as consistent as the prompt is specific

This is exactly the problem the character cultivation studio is built around: it keeps the appearance block fixed between rolls so only the outfit and scene change.

For users, the tell is the gallery. If a platform shows a grid of one character looking like one person across scenes, it is doing more than prompting. If the character looks like a different actor each time, it is not.

Route 3: video

Video is the newest and the least settled. Two routes:

The constraints are structural rather than temporary:

Expect video to have the lowest quality-per-credit of the three routes for some time, and to be the feature most often "included" in a plan in name only.

Why quotas are structured the way they are

Because the three routes have different marginal costs, platforms meter them differently, and the metering is usually the real content of the subscription.

Two things to check that the marketing page will not tell you:

  1. What happens at the limit. Does the feature stop, degrade, or silently switch to a different (cheaper) model? The last one is common and rarely stated up front.
  2. Whether credits expire. Monthly allowances that do not roll over are effectively a smaller number than they look.

Our pricing breakdown covers how the subscription, credit and per-generation models differ, and the terms checklist turns the checks into a printable list to run before paying.

What to test before subscribing for media

Key takeaways

For the vocabulary, see the glossary; for how media restrictions interact with policy layers, see content filters.