To add captions to a video on iPhone, record clean audio, generate a caption track from speech recognition, then correct the words that matter and style them for readability. Burned-in captions are the safer choice for TikTok, Reels, and Shorts, because they survive re-uploads and look the same on every account.
Captions are no longer a nice extra. A large share of short-form video is watched with the sound off or in a place where sound is rude, and a video without words on screen loses those viewers in the first second. The good news is that the whole job takes about five minutes once you stop doing it by hand.
Burned-in captions or platform captions
There are two kinds of captions and they behave differently.
Platform captions are generated by TikTok, Instagram, or YouTube after you upload. They are easy, they can be toggled off by the viewer, and their position and styling are decided by the app, not by you. Burned-in captions are drawn into the video pixels before upload. They cannot be switched off, they look identical everywhere, and they survive being downloaded and re-posted.
For UGC deliverables, burned-in captions are usually what a brand expects, because the client wants the file to look the same wherever it ends up. A useful compromise is to burn in your styled captions and still let the platform add its own text track for accessibility.
Start with audio the recognizer can read
Every caption workflow, automatic or manual, is limited by the recording. Speech recognition is good at clear speech in a quiet room and poor at everything else, so ten seconds of care before the take saves you five minutes of correction after it.
- Record indoors, away from a fan, an extractor, or an open window.
- Keep the microphone within about an arm's length of your mouth.
- Finish your words. Trailing off at the end of a sentence is the most common source of dropped text.
- Say brand names slightly slower than the rest of the sentence.
- Avoid speaking over music you plan to add later.
If you record with a teleprompter, this is easier than it sounds, because you are not searching for words while you talk. Our guide on using a teleprompter on iPhone covers the delivery habits that also happen to produce clean audio.
Generate the caption track
Once you have the take, you need timed text. Speech recognition produces a list of words with a start and end time for each one, and that timing is what lets captions appear in sync rather than as a static block.
Word-level timing is the detail to check. Some tools give you one caption per sentence, which is fine but flat. Word-level timing lets a caption highlight the word being spoken, which is what makes the caption feel alive instead of pasted on. UGCut generates captions with on-device speech recognition and a word-by-word active highlight, so the timing comes out of the same pass that produced the transcript.
If your editor also cuts your take into sentences, the caption boundaries and the cut boundaries agree by construction. That is the practical advantage of a script-anchored workflow, which we cover in auto-cutting video by sentence.
Fix only the words that change meaning
Do not proofread a caption track like a document. Read it once, looking for four things: the brand name, product names, numbers, and any word that reverses the meaning of a sentence. Fix those. Leave the rest.
A missing comma costs you nothing. A brand name spelled wrong in a paid deliverable can cost you the client. If the same name is wrong every time, fix the first instance and then check whether your tool can apply the same correction to the rest.
Break lines the way people speak
Caption line breaks are a readability decision, not a formatting accident. The reader should be able to take in a caption in one glance without moving their eyes across the whole screen.
- Keep each caption to a few words, roughly what you would say in one breath.
- Break at a natural pause, not mid-phrase.
- Never split a brand name or a number across two captions.
- Keep a caption on screen long enough to read at least once.
- Do not stack more than two lines at a time on a portrait video.
A quick test: mute your video and read only the captions. If you can follow the argument, the breaks are working. If you find yourself re-reading, the captions are too long or they change too fast.
Style captions for a small screen
Captions are read on a phone held at arm's length, often in daylight. Contrast beats decoration every time.
- Use a heavy weight rather than a thin one.
- Add a solid shade behind the text or a dark outline, so the words hold up over a bright background.
- Pick one style and keep it across a series, so your videos are recognizable at a glance.
- Reserve colour for one emphasised word, not for every second word.
- Check the caption over your busiest frame, not your simplest one.
Animated styles are effective in moderation. Movement attracts attention, which is useful once per sentence and exhausting for thirty seconds straight.
Keep captions inside the safe zone
Every platform draws its own interface over your video: a username, a caption, a row of buttons down one side, and a progress bar along the bottom. Text placed under that interface is simply invisible, and you will not notice in your editor because the interface is not there yet.
Keep captions in the middle band of the frame, clear of the top and bottom. A preview that shows the platform overlay saves you from finding out after publishing. UGCut's export presets include a safe-zone preview for TikTok, Reels, and Shorts for that reason.
Timing, and the word that carries the point
Good captions are slightly ahead of nothing and slightly behind nothing. If a caption appears before the word is spoken, viewers read ahead and stop listening. If it lags, the video feels broken.
Check sync at two places rather than across the whole video: the first three seconds, where you either keep the viewer or lose them, and the call to action at the end, where a mistimed caption undercuts the ask. If both are right, the middle is almost always right too.
Captions and accessibility are not the same job
Burned-in captions help sound-off viewers and they help many viewers with hearing loss, but they are not a full accessibility solution. Screen readers cannot read pixels, and a viewer who needs a different text size cannot change yours.
If accessibility matters for a deliverable, do both: burn in your styled captions for the sound-off audience, and also provide the platform caption track or a transcript. Writing a short description of what happens visually in the post text costs nothing and helps people who cannot see the screen.
On-device or uploaded
Some caption tools send your footage to a server to transcribe it. Others do the recognition on the phone. For a personal vlog the difference is minor. For an unreleased product under a non-disclosure agreement, it is the whole question.
UGCut runs its captioning on-device as part of the core loop, and the specifics of what stays on your phone are set out on the privacy page. Whatever tool you use, find out where the footage goes before you paste in a client brief.
A caption pass that takes five minutes
- Generate the caption track from the recorded take.
- Read it once and fix names, numbers, and reversed meanings.
- Check the line breaks by muting the video and reading along.
- Apply one style, then look at it over your busiest frame.
- Preview with the platform overlay and move anything hidden behind it.
Run the same five steps every time and captions stop being the part of the edit you dread.
Common caption problems
The captions are correct but hard to read
Increase the weight, add a dark outline or a solid shade behind the text, and reduce the number of words per caption. Readability is almost always a contrast and length problem rather than a font choice.
Captions disappear behind the TikTok interface
Move them toward the middle of the frame and preview with a safe-zone overlay. The bottom third of a portrait video belongs to the platform, not to you.
The recognizer keeps mangling one word
It is usually a proper noun or an unusual product name. Fix it once and apply the correction throughout, and say it a little more slowly in the next take.
Captions are out of sync after editing
This happens when captions are generated before the cuts are final. Cut first, then caption, or use a tool where captions and cuts come from the same transcript so they move together.
Captions are the cheapest attention you will ever buy. Get the audio right, let a recognizer do the typing, then spend your five minutes on the four or five words that actually carry the message.