How ClipWith reads your video
What the editor actually knows about your footage, and why that changes what you can ask for.
Most editing tools operate on a timeline of anonymous rectangles. ClipWith looks at the footage first, which is why you can refer to what's in the video rather than to timestamps.
What it has to work with
The words. Transcribed at the word level, with timings. This is why "cut the bit where I mention pricing" works, and why cuts land between words rather than through them.
The picture. On paid plans the video is indexed — objects, faces, actions, scene changes. This is what lets "cut to the shot where I hold up the product" resolve to an actual moment.
Real frames. When you ask for something, actual stills from your video go to the model alongside your request. It isn't reasoning about a description of your video; it's looking at it.
Where people are. Face and body positions per frame, which is what caption placement and reframing depend on.
Why this matters for what you type
Because it can see, you can be vague in useful ways:
cut the boring part
punch in when I make a point
put the captions where they won't cover my face
Each of those requires knowing what's on screen. A tool working from timestamps alone would need you to specify them.
What it does after editing
It renders frames from the result and looks at those too. If a caption collided with a face or a cut landed mid-word, it corrects before handing back. That check is why an edit takes a few seconds rather than being instant.
Free plans
Free accounts skip visual indexing — cost, not policy. Transcript-based edits still work. Anything that depends on seeing the picture will tell you it needs a paid plan rather than guessing.