Auto captions get words wrong when the recognizer has to guess: unusual brand names, numbers spoken quickly, background noise, overlapping speech, and sentences that trail off. Fix the mistakes that change meaning, leave the rest alone, and change how you record so the same word stops failing on the next take.
Automatic captions are close enough to correct that the errors are surprising when they appear, and they always appear on the words you most needed to be right. Once you know why a recognizer fails, most of the errors become preventable at the recording stage rather than repairable at the editing stage.
What speech recognition is actually doing
A speech recognizer does not hear words. It receives sound, breaks it into small slices, and produces the sequence of words most likely to have produced that sound, weighted by what usually follows what in the language it was trained on. That second half is why errors look the way they do.
The model prefers common phrasing. When your audio is ambiguous, it falls back on what is statistically likely, which is why a rare product name becomes an ordinary word that sounds a bit like it. It is not confused. It is guessing, and it is guessing in the direction of the ordinary.
Brand names and invented words
This is the single most common failure in UGC captions, and the most expensive, because the brand name is the one word the client will check.
Invented names, deliberate misspellings, and names that combine two words are all outside the recognizer's normal vocabulary, so it substitutes something familiar. Nothing you do in the recording will make an unknown word known, but you can make the audio unambiguous.
- Say the name a little slower than the words around it.
- Put a small pause before it rather than running it into the previous word.
- Give it its own sentence in the script the first time it appears.
- Avoid saying it over music or while handling packaging.
- Check that first instance before anything else in the caption pass.
Numbers, prices, and units
Numbers fail in two ways. The digits can be wrong, and the formatting can be wrong in a way that changes the meaning: fifteen becomes fifty, a percentage becomes a plain number, or a price loses its currency.
Spoken numbers are short and often unstressed, which gives the recognizer very little sound to work with. If a discount code or a price appears in a brief, treat it like a brand name: slow down, separate it from the surrounding words, and verify it by eye afterwards. This is one of the few places where a caption error can create a genuine problem for the client rather than an aesthetic one.
Homophones the model cannot resolve
Some errors are not fixable at the audio level, because the words genuinely sound the same. Whether the recognizer writes "their" or "there" depends on context, and short-form video gives it very little context to use.
These errors are usually harmless to comprehension, and viewers reading captions at speed rarely notice them. Fix them if the caption is prominent and the video is a paid deliverable. Otherwise, let them go and spend the time on something the viewer will actually notice.
Noise, music, and two people talking
Background sound is the enemy of transcription, and some kinds are worse than others. Steady noise like a fan or traffic raises the floor and blurs consonants. Sudden noise, a door or a notification, can take out a whole word. Music with vocals is the worst case, because the recognizer will try to transcribe the singer.
Two people talking at once produces the strangest output, because the model tries to build one sentence from two. If you shoot with a second person, take turns cleanly and leave a small gap between speakers rather than overlapping for energy.
Add music after the captions are generated, never before. If you already have a mixed track, expect to correct more than usual.
Sentences that trail off
Most people lower their volume across a sentence and drop the last word almost entirely. You know what you said, so you do not hear the gap. The recognizer does hear it, and either omits the word or replaces it with a shorter one.
The fix is delivery, not software. Finish the last word of every sentence at the same volume as the first. If you use a teleprompter, this is easier, because you are not simultaneously trying to remember the next line. The pacing habits in our guide on using a teleprompter on iPhone are the same habits that produce clean transcripts.
Accents and speaking speed
Recognition quality varies with accent, and there is no polite way around that. If your accent is under-represented in the training data, you will see a higher error rate than a creator with a different one, on the same equipment and in the same room.
What you can control is speed and articulation. Speaking slightly slower than your natural conversational pace, and finishing consonants at the ends of words, improves recognition more than any setting in an app. Check the language and region setting on your device too, since a mismatch between your spoken variety and the selected one produces avoidable spelling differences.
Which errors to fix and which to ignore
Correcting a transcript word by word is a bad use of an hour. Sort the errors into three buckets and only act on the first two.
- Fix anything that changes meaning: a reversed negative, the wrong number, a substituted claim.
- Fix anything a client will check: the brand name, product names, the discount code, the required phrase from the brief.
- Ignore punctuation, capitalisation, and homophones that read fine at speed.
Viewers read captions in fragments while listening to you speak. Small textual imperfections vanish in that mode. A wrong price does not.
Fix it once, then apply it everywhere
If a word is wrong in one caption, it is usually wrong in every caption. Rather than editing each instance, correct the first one and look for a way to apply the same correction across the transcript.
This is also where a script-anchored workflow saves time. When captions come from the same transcript that produced your cuts, a correction lands in one place instead of being re-entered per clip. The mechanics of that anchoring are described in auto-cutting video by sentence, and the caption side of it is covered in how to add captions to a video on iPhone.
UGCut generates its captions with on-device speech recognition in the same pass that produces the sentence cuts, so the words on screen and the cut points come from one transcript rather than two.
Recording habits that prevent the same error
Every correction you make is a note about your next shoot. The list below is short because it is the list that actually changes error rates.
- Record in the quietest room you have, with the door closed.
- Keep the microphone within about an arm's length.
- Turn off anything that makes a steady hum before you start.
- Finish your sentences at full volume.
- Give brand names and numbers their own space in the script.
- Record a ten-second test and read its transcript before the real take.
That last one is the highest-value habit on the list. Ten seconds of test audio tells you whether today's room, today's microphone, and today's brand name are going to cooperate, while you can still change all three.
When to re-record instead of correcting
Sometimes the transcript is telling you the audio is bad rather than the recognizer is bad. If a whole sentence comes back as nonsense, do not fix the text. Fix the sound and say the line again.
Re-recording one sentence should be cheap. If your workflow keeps takes per script line, you can replace a single failed sentence without touching the rest of the video, which changes the calculation entirely: a two-second retake beats five minutes of typing. UGCut's takes review is built around that per-sentence unit.
Common questions about automatic caption errors
Can I teach the recognizer my brand name?
Some tools accept a custom vocabulary or let you save a replacement rule. Where that is not available, the practical equivalent is correcting the first instance and applying it across the transcript, then saying the name more deliberately next time.
Does a better microphone fix caption errors?
It helps, mostly by improving the ratio between your voice and the room. It does not help with unknown words, homophones, or numbers spoken too fast, so treat it as one improvement among several rather than a solution.
Is on-device recognition less accurate than a server?
Accuracy depends on the model, the language, and your audio rather than on where the computation happens. The clearer trade-off is what leaves your phone: on-device processing keeps unreleased product names and client scripts local. UGCut's privacy page sets out what stays on the device.
Should I just type the captions myself?
Only for very short videos. Typing and timing captions by hand for a sixty-second video takes far longer than correcting five words in a generated transcript, and hand-typed captions still need the same styling and safe-zone work afterwards.
Treat auto captions as a first draft that is ninety percent right. Your job is not to proofread it. Your job is to find the five words that matter, fix them, and change one thing about how you record so those five words behave next time.