VERTICAL CAPTION GUIDE

Design captions for the screen people actually watch on.

A 9:16 canvas is not entirely usable space. Platform controls, account details, descriptions, and interaction buttons compete with faces, demonstrations, and captions, so crop and text need to be planned together.

A creator reviewing vertical video and caption placement on screen
Caption quality is judged on moving footage at phone size, not on an isolated desktop frame.

A durable caption system combines conservative placement, phone-readable type, natural phrase groups, strong contrast, faithful wording, synchronized timing, and final-file review.

01

Understand why the vertical frame has less usable space

Short-form interfaces often cover the top, bottom, and right side of a video with navigation, profile information, descriptions, audio labels, progress indicators, and interaction controls. The exact overlays differ by platform and can change over time.

Treat safe areas as conservative working zones rather than universal pixel guarantees. Keep essential text closer to the central viewing area and test the exported clip in the actual destination when possible.

  • Avoid extreme edges
  • Reserve interface space
  • Test real destinations
02

Choose caption placement together with the crop

Vertical reframing and captions compete for the same limited canvas. A centered talking head may leave room below the face, while a screen recording may require text above a key control. A two-person interview can change available space whenever the camera angle switches.

Establish the visual priority first: active speaker, essential object, chart, demonstration, or interface region. Then place captions where they remain readable without covering that priority through the entire segment.

  • Visual priority first
  • Caption zone second
  • Every scene change checked
03

Use a type size designed for a real phone

Captions are viewed at arm’s length on a small screen, often over moving footage. Text that feels elegant on a desktop preview may become unreadable at normal mobile size. Evaluate scale on a phone-sized viewport without zooming the editor.

Choose clear letterforms, enough weight to survive bright and dark backgrounds, and spacing that does not collapse. Avoid making captions so large that normal phrases break into a rapid series of one- or two-word lines.

  • Phone-first scale
  • Clear letterforms
  • Balanced line length
04

Group words into phrases the viewer can scan

Raw transcripts follow speech rather than visual reading. Break captions at natural phrase boundaries so each group communicates a small unit of meaning. Keep names, numbers, technical terms, and closely related words together.

Dense verbatim blocks force the viewer to choose between reading and watching. Constant word-by-word animation creates the opposite problem: too much motion. Stable one- or two-line phrase groups usually provide a calmer reading rhythm.

  • Natural phrase boundaries
  • Limited line count
  • Stable text changes
05

Create contrast that survives changing footage

A caption may sit over a dark jacket in one frame and a bright wall in the next. Color chosen from a still image is not enough. Use a foreground and restrained outline, shadow, or backing treatment that remains legible across the complete clip.

Check the busiest and brightest frames, not only the opening. Accent color can emphasize a short phrase, but it should not reduce readability or imply meaning that the speaker did not express.

  • Bright and dark frames tested
  • Consistent treatment
  • Restrained accent
06

Edit speech into clean-verbatim text

Spoken language contains filler, false starts, repetition, and self-correction. Clean-verbatim editing can remove distracting artifacts while preserving the claim, tone, order, and meaningful uncertainty. It is a readability process, not permission to rewrite the speaker.

Do not strengthen tentative language, remove a qualification, invent a smoother conclusion, or turn an interview answer into promotional copy. Verify names, brands, figures, and unfamiliar terms against the source audio.

  • Noise removed, not meaning
  • Qualifications retained
  • High-risk terms verified
07

Synchronize caption changes with audible phrases

Text that arrives too early can reveal the payoff; text that lags forces the viewer to read one phrase while hearing the next. Both problems become more obvious at natural playback speed than on a paused timeline.

Use timestamps as the starting point, then watch the rendered clip. Breaths, pauses, overlap, and rapid delivery may require adjustments. Keep picture and source audio trimmed at matching boundaries so captions are not compensating for a synchronization error.

  • Phrase-level timing
  • No premature payoff
  • Audio-video sync confirmed
08

Test sound-on and sound-off viewing

On sound-off playback, captions need enough context to carry the central idea. On sound-on playback, they should reinforce comprehension without competing with facial expression, demonstrations, or pacing.

If the clip makes no sense without sound because a key phrase is missing, correct the text. If captions dominate when sound is on, reduce density, motion, or visual weight. The same export should remain coherent in either mode.

  • Meaning survives without sound
  • Text supports the voice
  • Attention remains balanced
09

Verify the exported MP4 and document exceptions

Font rendering, line wrapping, crop, scale, and timing can change during export. Watch the final 9:16 file from first frame to last at phone size. Look for clipped letters, hidden text, captions over faces, phrases that disappear too quickly, and text that outlives the audio.

A reusable system can define typeface, weight, base size, maximum width, line count, safe zone, contrast treatment, and accent behavior. Document exceptions for screen recordings, moving speakers, dense lower thirds, or unusually long technical phrases instead of abandoning the system entirely.

  • Full export playback
  • Reusable base rules
  • Exceptions reviewed deliberately

CAPTION QUESTIONS

What should be checked before vertical captions are approved?

Is there one safe area that works on every platform?

No. Interfaces differ and change. Use a conservative central zone, protect the important visual, and test the final export in each destination when possible.

How many caption lines should appear at once?

Use as few as needed for a readable phrase, commonly one or two. Avoid dense blocks that make viewers choose between reading and watching.

Should captions animate word by word?

Stable phrase groups are usually calmer and easier to verify. Motion should have a clear reading benefit rather than becoming the main visual event.

Why check the exported MP4?

Export can change font rendering, wrapping, crop, scale, and timing. Only the final file proves what viewers will actually receive.

Ready to see captions on a real rendered clip?

Create a job, watch the 9:16 output at phone size, and verify wording, placement, timing, and source audio together.

Create your clips