Captions

Captions that know where the face is

Burned-in captions placed around your subject, timed to the word, in a look you can change by asking.

A talking-head frame with captions sitting across the speaker's chin
Fixed-position captions, over the face
the editor saw

speaker centred, head fills upper third

→ captions moved below chest line, 2 words per card

The same frame with captions placed below the speaker's chest
Placed clear of the subject

How it decides

It finds the speaker first

Every frame is checked for where the face and body actually are, so the caption block lands in empty space rather than a fixed percentage down the frame. When the subject moves, the captions follow.

Timing comes from the words, not the line

Transcription is word-level, so each word lights on the syllable it belongs to. Cards break on natural phrase boundaries instead of a character count.

Emphasis is a typography decision

The word carrying the point gets the treatment — a heavier face, a colour flip, a size step. You can set which words, or let it choose from the sentence's own stress.

What you control

Things you can type. The editor answers to language, not menus.

  • make the captions bigger and put them at the top
  • use the mono look, all caps, no emoji
  • highlight every product name in orange
  • two words at a time, tighter timing

What it won't do

  • It won't invent words the speaker didn't say — captions come from the audio, and a mis-heard word is a transcription fix, not a rewrite.
  • It won't caption over a face when there's nowhere clear to go; it shrinks the block instead of covering the subject.