Getting True 9:16 Vertical Video Out of an AI Model
Aspect ratio is a generation parameter, not a prompt word — writing "vertical" or "9:16" in the text does nothing if the request isn’t sending an aspect value, and the model falls back to landscape. Set it at generation time and target 1080×1920. Cropping a landscape render to vertical discards about 60% of a frame the model deliberately composed for a wide screen, usually including the part you needed.
The parameter, not the prompt
This one catches almost everybody once. You write a careful prompt that says "vertical 9:16 format, portrait orientation," and a landscape clip comes back. The prompt text and the frame shape travel on different channels: the text conditions what’s in the picture, while the picture’s dimensions come from a separate aspect parameter on the request.
When that parameter is absent, unsupported for the model in question, or set to a value the provider rejects, you get the default — which is landscape for nearly every model, because that’s what the training corpus is. The failure is silent, which is why people reasonably conclude the model is ignoring them.
Worth checking. Aspect support varies per model, and some accept a different set of ratios than their own documentation implies. If vertical works on one model in a tool and not another, that’s a routing problem, not your prompt.
Why cropping isn’t the shortcut it looks like
The tempting fix — generate landscape, crop to vertical — fails for a reason that isn’t about pixels. A model composing a 16:9 frame stages for 16:9: it puts subjects off-centre, leaves negative space at the sides, gives a wide shot the headroom a wide shot wants, and places meaningful detail across the full width.
A centre crop to 9:16 keeps about 40% of the width and throws the staging away. Two people talking become one person and an elbow. The object the shot existed to show is outside the crop. And you can’t recover it in the edit, because the information was never in the centre ninth.
Vertical is a different grammar
Getting the frame shape right is necessary and not sufficient — vertical shots want different coverage:
- Go closer. Medium close-up is the workhorse size. There isn’t horizontal room for a comfortable two-shot.
- Singles over two-shots. Cover conversation in alternating singles; it’s clearer and it cuts better.
- Cut faster. Vertical drama runs shorter average shot lengths than landscape — often two to four seconds.
- Use the height. Stairwells, standing over someone seated, low angles up — vertical composition rewards vertical geometry.
- Wides do less. A vertical wide establishes very little, so don’t spend the time on it that you would in landscape.
Safe areas
The delivered frame is not the visible frame. Platform interface — captions, handles, buttons, the description — covers the bottom fifth and the right edge. Anything staged there is obscured in the only place the video will actually be watched.
Keep faces and subject matter in the middle band, put burnt-in captions above the bottom fifth rather than at the base of the frame, and never rely on detail near an edge. Design for what the viewer sees, not for what the file contains.
Don’t pad to fit
Letterboxing or pillarboxing a mismatched render to reach 1080×1920 is worth avoiding entirely. Bars consume the safe area, they read as recycled landscape content, and platform algorithms treat them as a quality signal in the direction you don’t want. If a shot came back the wrong shape, regenerate it at the right shape.
How ShowMaker Studio handles it
Format is a project-level decision, so a vertical project renders vertical everywhere — stills, video, and the final assembly all inherit 9:16 without per-shot configuration, and model routing accounts for which models support the aspect natively. Coverage generated for a vertical project favours the closer sizes vertical needs, and captions are placed inside the safe band by default rather than at the bottom of the frame. The micro drama format is built around this end to end, and the publishing guide covers export targets for TikTok, Reels, and Shorts.
Related questions
Why does my AI video come out 16:9 when I asked for 9:16?
Because aspect ratio is a generation parameter, not a description. Writing "vertical" or "9:16" in the prompt text usually does nothing — the model renders at whatever its aspect parameter says, and if that parameter is missing or invalid it falls back to its landscape default. Check that the request is actually sending an aspect value, not that the words appear in the prompt.
Can I just crop a 16:9 AI video to 9:16?
You can, but you throw away the composition. A model given a landscape frame stages the action for landscape — subjects at the sides, headroom for a wide frame, important detail outside the centre ninth. Cropping to vertical cuts through that staging, so you lose roughly 60% of the image and often the part that mattered. Generate vertical instead.
Why is my subject's head cut off in vertical AI video?
Vertical framing needs different staging, and most models are trained predominantly on landscape footage. Ask explicitly for the framing you want in vertical terms — medium close-up, subject centred, head-to-waist in frame, headroom at the top — rather than assuming the model will adapt a landscape instinct to a tall frame.
What resolution should vertical AI video be?
1080×1920 is the standard target for TikTok, Reels, and Shorts. Generate at the highest vertical resolution your model supports and downscale to that if needed — downscaling is clean, upscaling is not. Never letterbox or pillarbox to reach the frame size; platforms treat bars as low-quality and they eat the safe area.
Do captions and UI safe areas matter in vertical video?
A great deal. Platform interface elements cover the bottom fifth and the right edge of the screen, so anything staged there — captions included — gets obscured. Keep captions and any important visual information inside the middle band, and don't compose a shot whose subject sits in the bottom sixth of the frame.
Should vertical shots be framed differently from landscape ones?
Yes, meaningfully. Vertical rewards closer sizes: singles over two-shots, medium close-ups over mediums, and faces over environments. There's very little horizontal room to establish geography, so wides do less work and cuts happen faster. Micro-drama averages notably shorter shot lengths than landscape drama for exactly this reason.
Every fix on this page is a feature, not a workaround.
ShowMaker Studio is built around these problems — reference-locked casting, location look signatures, coverage-based scenes, delivery-timed captions, and per-shot model routing. Three free films a month, no card.