Why Your AI Video Captions Drift Out of Sync
Captions that start on time and finish seconds late were timed against the wrong clock — the script text or an estimated duration, rather than the audio that actually survived into the final mix. Estimation error accumulates line by line. The fix is to generate caption timings by transcribing the delivered audio, after all trims, speed ramps, and edits, and to burn them last, after every overlay.
Growing error tells you the cause
The shape of the drift is diagnostic. A constant offset — everything a beat late from the first word — is a start-time problem, usually an audio track laid down at the wrong position. That’s easy and rare.
Error that grows is a different animal. It means each line’s start time is computed from the end of the previous line, and those ends are estimates. Estimated speech duration is a decent approximation — around two and a half words a second — but a performed line is never exactly its estimate. It has a pause before the important word, a breath, an emphasis that stretches a syllable. Every one of those differences is added to the running total, and by line twenty you’re seconds out.
Time captions to what survived, not to what you wrote
The script is intent. The audio is fact, and there are several steps between them where they diverge:
- The voice performance is slower or faster than estimated, line by line.
- The take says something slightly different from the script — a contraction, a dropped word.
- The clip was trimmed, moving everything after the trim point.
- A speed ramp compressed part of the shot, moving words within it non-linearly.
- The mix ducked, faded, or replaced part of the track.
Only the finished audio reflects all five. Transcribing that audio with word-level timestamps gives you captions that are correct by construction — there is no estimate left to accumulate.
The subtle version. Where a clip has both a native-audio take and a separate voiceover, captions must be timed to whichever one is in the final mix. Timing against the take that got muted produces captions that are internally consistent and completely wrong.
Order of operations
Captions are the last thing you burn. The sequence that holds up:
- Assemble picture and cut it. All trims, all reorders, all speed ramps. Lock the edit.
- Mix the audio. Dialogue, score, effects, ducking. This is the track captions will be derived from.
- Composite every overlay. Cutaways, B-roll, title cards, name tags, watermarks.
- Transcribe the mix and burn captions on top. Word-level timings from the actual audio, painted over the finished composite.
Get this order wrong and you produce the second most common caption bug: text that is perfectly synced and invisible, because a cutaway was painted over it.
Why this matters more for vertical
Short-form vertical video is largely watched with the sound off. Captions aren’t an accessibility checkbox there — they are the primary channel for the dialogue, and they are doing retention work in the first three seconds. A caption a beat behind the cut reads as low quality to an audience that would never articulate why.
How ShowMaker Studio handles it
Captions are transcribed from the audio that actually survives the compose step, with word-level timings, and recalculated when a clip is trimmed or speed-ramped. They’re burnt after cutaways and overlays rather than before, so nothing paints over them. Styling — font, lower-third placement, per-word hold, slide size — is directorial rather than fixed, and on-screen text nobody speaks can be captioned explicitly. The editor guide covers the caption controls.
Related questions
Why do my AI video captions drift out of sync?
Almost always because they were timed against the wrong clock — the script text or an estimated duration rather than the audio that actually ended up in the finished mix. Any difference between the estimate and the real delivery accumulates line by line, so the captions start close and end seconds adrift. Timing them from the delivered audio removes the drift entirely.
Why are my captions correct at the start and wrong at the end?
That's the signature of accumulated error. A fixed offset would be wrong from the first word; growing error means each line's timing is derived from the previous line's estimated length. It's a strong indicator that captions were generated from text rather than transcribed from the finished audio.
Should captions be generated from the script or from the audio?
From the audio. The script is what you intended; the audio is what happened. Performance pauses, emphasis, breaths, retakes, and speed ramps all change timing, and only a transcription of the final mix reflects them. Transcribing also catches the case where the delivered take says something slightly different from the script.
Why do captions break after I trim or speed up a clip?
Because the caption timings referenced the original clip, and trimming or ramping moved every word after the edit point. Captions have to be recalculated against the post-edit timeline, not the pre-edit one. Any pipeline that bakes captions before the final trim will drift the moment you touch the cut.
Why did my captions disappear under a cutaway or B-roll?
Because the overlay was composited on top of already-burnt-in captions. If captions are burnt into a base clip and a cutaway is painted over that clip, the cutaway covers the text. Captions should be burnt last, after every overlay, or re-burnt over the composite.
Do burnt-in captions matter for vertical video reach?
Substantially. A large share of short-form vertical video is watched muted, so captions are often the only way the dialogue lands at all. On TikTok, Reels, and Shorts they function as a retention mechanism rather than an accessibility afterthought — which is why sync accuracy on those formats is worth real attention.
Every fix on this page is a feature, not a workaround.
ShowMaker Studio is built around these problems — reference-locked casting, location look signatures, coverage-based scenes, delivery-timed captions, and per-shot model routing. Three free films a month, no card.