Voice, Image and Video in AI Companions: Three Technical Routes
A companion app is usually described as one product. Technically it is a chat system with up to three unrelated generation pipelines attached: speech, images and video. Each is built on different technology, fails in different ways, and costs the platform different amounts of money. That is why the quotas and the quality vary so much between them.
One product, four systems
The text chat is a language model producing tokens. The other three are entirely separate:
- Speech turns text into audio, and sometimes audio into a cloned voice.
- Images turn text into a picture, and increasingly a reference picture into a consistent picture of the same character.
- Video turns text or an image into a short moving clip.
They do not share code paths, they are not the same model, and on most platforms they do not share quota. When a subscription "includes voice, images and video", the useful question is which of the three is actually usable within the monthly allowance.
Route 1: speech
Two things get called voice, and they are different products.
Text-to-speech converts the character's reply into audio. Modern neural TTS is good enough to be mistaken for a recording in short sentences. The engineering problems are latency and streaming: audio has to start before the whole reply is generated, or the character feels slow to answer.
Voice cloning takes a sample of a voice and makes the model speak in it. The technical bar has fallen far enough that a short sample is often sufficient, which is why consent and misuse are the live issues here rather than quality.
Practical characteristics:
- Latency is the whole experience. Text can arrive word by word and feel responsive. Audio has to be buffered, so the same reply feels slower. Products that hide this well are streaming the speech pipeline in parallel with the text model.
- Cost is per character of audio, not per message. A paragraph spoken aloud costs materially more than the same paragraph typed, which is why voice is usually metered separately.
- The failure mode is uncanny, not broken. Wrong emphasis, flat affect in emotional scenes, mispronounced names, and a voice that stays cheerful while the character is supposed to be upset.
If voice quality matters to you, test it on the specific thing you will use it for: reading a long emotional reply, not a sample sentence.
Route 2: images
Image generation in this category is almost entirely diffusion-based: the model starts from noise and iteratively removes it, guided by your text.
The interesting problem is not generating a picture. It is generating the same person twice.
A text prompt describes a character in general terms, and a diffusion model has no memory of the last picture it made. Ask for the same character in a different scene and you get a different face. Three approaches are used to fight this:
| Approach | How it works | Trade-off |
|---|---|---|
| Reference or style conditioning | The previous image, or a reference image, is fed in alongside the prompt | Simple; identity still drifts over repeated generations |
| Trained character adapters | A small model is trained on a handful of images of one character, then loaded at generation time | Much stronger consistency; needs a training step and per-character storage |
| Structured prompt discipline | A fixed appearance block is reused verbatim in every prompt | No extra cost; only as consistent as the prompt is specific |
This is exactly the problem the character cultivation studio is built around: it keeps the appearance block fixed between rolls so only the outfit and scene change.
For users, the tell is the gallery. If a platform shows a grid of one character looking like one person across scenes, it is doing more than prompting. If the character looks like a different actor each time, it is not.
Route 3: video
Video is the newest and the least settled. Two routes:
- Image to video. Generate a still image first, then animate it. Most companion-app "video" features work this way, because the still frame is where the character consistency problem was already solved.
- Text to video. Generate from a description directly. More flexible, harder to control, and further from being reliable for a specific character.
The constraints are structural rather than temporary:
- Temporal consistency. A still image only has to look right once. A clip has to stay coherent frame to frame, which is a much harder generation problem. Faces and hands flicker first.
- Compute cost per second. Video generation is orders of magnitude more expensive than text. This is why clips are short, resolutions are modest, and video almost always sits behind the tightest quota on the pricing page.
- Latency. A generation that takes long enough that the product cannot present it as a conversation turn. Video is usually an asynchronous request, not a reply.
Expect video to have the lowest quality-per-credit of the three routes for some time, and to be the feature most often "included" in a plan in name only.
Why quotas are structured the way they are
Because the three routes have different marginal costs, platforms meter them differently, and the metering is usually the real content of the subscription.
- Text is metered by message or token budget, sometimes with a "memory size" tier on top.
- Images are metered per generation or in credits, with resolution tiers.
- Video is metered per clip, with a tight monthly cap.
- Voice is metered by character count or minutes of audio.
Two things to check that the marketing page will not tell you:
- What happens at the limit. Does the feature stop, degrade, or silently switch to a different (cheaper) model? The last one is common and rarely stated up front.
- Whether credits expire. Monthly allowances that do not roll over are effectively a smaller number than they look.
Our pricing breakdown covers how the subscription, credit and per-generation models differ, and the terms checklist turns the checks into a printable list to run before paying.
What to test before subscribing for media
- Voice: a long, emotionally varied reply, not a greeting. Listen for emphasis and for whether the voice changes tone with the text.
- Images: the same character in three different scenes. Compare the faces side by side; that is the only test that matters.
- Video: one clip, watched twice. Look at hands and the background between frames.
- Quotas: spend one month's allowance in an afternoon and see what it actually bought, then check what happens when you hit the cap.
Key takeaways
- Speech, image and video are separate pipelines with separate models, costs and failure modes — not features of one system.
- Voice is limited by latency and per-character cost, and its typical failure is uncanny delivery rather than broken output.
- Image generation's hard problem is identity consistency, solved by reference conditioning, trained character adapters, or strict prompt discipline.
- Video is structurally expensive: clips are short and quotas are tight because compute cost scales with seconds.
- Media is metered separately from text, so the value of a subscription depends on which of the four systems you actually use.
- Test media the way you will use it — long replies, the same character across scenes, a clip watched twice — rather than with sample outputs.
For the vocabulary, see the glossary; for how media restrictions interact with policy layers, see content filters.