Captions
Captions that know where the face is
Burned-in captions placed around your subject, timed to the word, in a look you can change by asking.
speaker centred, head fills upper third
→ captions moved below chest line, 2 words per card
How it decides
It finds the speaker first
Every frame is checked for where the face and body actually are, so the caption block lands in empty space rather than a fixed percentage down the frame. When the subject moves, the captions follow.
Timing comes from the words, not the line
Transcription is word-level, so each word lights on the syllable it belongs to. Cards break on natural phrase boundaries instead of a character count.
Emphasis is a typography decision
The word carrying the point gets the treatment — a heavier face, a colour flip, a size step. You can set which words, or let it choose from the sentence's own stress.
What you control
Things you can type. The editor answers to language, not menus.
- make the captions bigger and put them at the top
- use the mono look, all caps, no emoji
- highlight every product name in orange
- two words at a time, tighter timing
What it won't do
- It won't invent words the speaker didn't say — captions come from the audio, and a mis-heard word is a transcription fix, not a rewrite.
- It won't caption over a face when there's nowhere clear to go; it shrinks the block instead of covering the subject.

