Why Auto Captions Get Words Wrong, and How to Fix Them Fast

The UGCut Team8 min read

Auto captions get words wrong when the recognizer has to guess: unusual brand names, numbers spoken quickly, background noise, overlapping speech, and sentences that trail off. Fix the mistakes that change meaning, leave the rest alone, and change how you record so the same word stops failing on the next take.

Automatic captions are close enough to correct that the errors are surprising when they appear, and they always appear on the words you most needed to be right. Once you know why a recognizer fails, most of the errors become preventable at the recording stage rather than repairable at the editing stage.

What speech recognition is actually doing

A speech recognizer does not hear words. It receives sound, breaks it into small slices, and produces the sequence of words most likely to have produced that sound, weighted by what usually follows what in the language it was trained on. That second half is why errors look the way they do.

The model prefers common phrasing. When your audio is ambiguous, it falls back on what is statistically likely, which is why a rare product name becomes an ordinary word that sounds a bit like it. It is not confused. It is guessing, and it is guessing in the direction of the ordinary.

Brand names and invented words

This is the single most common failure in UGC captions, and the most expensive, because the brand name is the one word the client will check.

Invented names, deliberate misspellings, and names that combine two words are all outside the recognizer's normal vocabulary, so it substitutes something familiar. Nothing you do in the recording will make an unknown word known, but you can make the audio unambiguous.

  • Say the name a little slower than the words around it.
  • Put a small pause before it rather than running it into the previous word.
  • Give it its own sentence in the script the first time it appears.
  • Avoid saying it over music or while handling packaging.
  • Check that first instance before anything else in the caption pass.

Numbers, prices, and units

Numbers fail in two ways. The digits can be wrong, and the formatting can be wrong in a way that changes the meaning: fifteen becomes fifty, a percentage becomes a plain number, or a price loses its currency.

Spoken numbers are short and often unstressed, which gives the recognizer very little sound to work with. If a discount code or a price appears in a brief, treat it like a brand name: slow down, separate it from the surrounding words, and verify it by eye afterwards. This is one of the few places where a caption error can create a genuine problem for the client rather than an aesthetic one.

Homophones the model cannot resolve

Some errors are not fixable at the audio level, because the words genuinely sound the same. Whether the recognizer writes "their" or "there" depends on context, and short-form video gives it very little context to use.

These errors are usually harmless to comprehension, and viewers reading captions at speed rarely notice them. Fix them if the caption is prominent and the video is a paid deliverable. Otherwise, let them go and spend the time on something the viewer will actually notice.

Noise, music, and two people talking

Background sound is the enemy of transcription, and some kinds are worse than others. Steady noise like a fan or traffic raises the floor and blurs consonants. Sudden noise, a door or a notification, can take out a whole word. Music with vocals is the worst case, because the recognizer will try to transcribe the singer.

Two people talking at once produces the strangest output, because the model tries to build one sentence from two. If you shoot with a second person, take turns cleanly and leave a small gap between speakers rather than overlapping for energy.

Add music after the captions are generated, never before. If you already have a mixed track, expect to correct more than usual.

Sentences that trail off

Most people lower their volume across a sentence and drop the last word almost entirely. You know what you said, so you do not hear the gap. The recognizer does hear it, and either omits the word or replaces it with a shorter one.

The fix is delivery, not software. Finish the last word of every sentence at the same volume as the first. If you use a teleprompter, this is easier, because you are not simultaneously trying to remember the next line. The pacing habits in our guide on using a teleprompter on iPhone are the same habits that produce clean transcripts.

Accents and speaking speed

Recognition quality varies with accent, and there is no polite way around that. If your accent is under-represented in the training data, you will see a higher error rate than a creator with a different one, on the same equipment and in the same room.

What you can control is speed and articulation. Speaking slightly slower than your natural conversational pace, and finishing consonants at the ends of words, improves recognition more than any setting in an app. Check the language and region setting on your device too, since a mismatch between your spoken variety and the selected one produces avoidable spelling differences.

Which errors to fix and which to ignore

Correcting a transcript word by word is a bad use of an hour. Sort the errors into three buckets and only act on the first two.

  1. Fix anything that changes meaning: a reversed negative, the wrong number, a substituted claim.
  2. Fix anything a client will check: the brand name, product names, the discount code, the required phrase from the brief.
  3. Ignore punctuation, capitalisation, and homophones that read fine at speed.

Viewers read captions in fragments while listening to you speak. Small textual imperfections vanish in that mode. A wrong price does not.

Fix it once, then apply it everywhere

If a word is wrong in one caption, it is usually wrong in every caption. Rather than editing each instance, correct the first one and look for a way to apply the same correction across the transcript.

This is also where a script-anchored workflow saves time. When captions come from the same transcript that produced your cuts, a correction lands in one place instead of being re-entered per clip. The mechanics of that anchoring are described in auto-cutting video by sentence, and the caption side of it is covered in how to add captions to a video on iPhone.

UGCut generates its captions with on-device speech recognition in the same pass that produces the sentence cuts, so the words on screen and the cut points come from one transcript rather than two.

Recording habits that prevent the same error

Every correction you make is a note about your next shoot. The list below is short because it is the list that actually changes error rates.

  • Record in the quietest room you have, with the door closed.
  • Keep the microphone within about an arm's length.
  • Turn off anything that makes a steady hum before you start.
  • Finish your sentences at full volume.
  • Give brand names and numbers their own space in the script.
  • Record a ten-second test and read its transcript before the real take.

That last one is the highest-value habit on the list. Ten seconds of test audio tells you whether today's room, today's microphone, and today's brand name are going to cooperate, while you can still change all three.

When to re-record instead of correcting

Sometimes the transcript is telling you the audio is bad rather than the recognizer is bad. If a whole sentence comes back as nonsense, do not fix the text. Fix the sound and say the line again.

Re-recording one sentence should be cheap. If your workflow keeps takes per script line, you can replace a single failed sentence without touching the rest of the video, which changes the calculation entirely: a two-second retake beats five minutes of typing. UGCut's takes review is built around that per-sentence unit.

Common questions about automatic caption errors

Can I teach the recognizer my brand name?

Some tools accept a custom vocabulary or let you save a replacement rule. Where that is not available, the practical equivalent is correcting the first instance and applying it across the transcript, then saying the name more deliberately next time.

Does a better microphone fix caption errors?

It helps, mostly by improving the ratio between your voice and the room. It does not help with unknown words, homophones, or numbers spoken too fast, so treat it as one improvement among several rather than a solution.

Is on-device recognition less accurate than a server?

Accuracy depends on the model, the language, and your audio rather than on where the computation happens. The clearer trade-off is what leaves your phone: on-device processing keeps unreleased product names and client scripts local. UGCut's privacy page sets out what stays on the device.

Should I just type the captions myself?

Only for very short videos. Typing and timing captions by hand for a sixty-second video takes far longer than correcting five words in a generated transcript, and hand-typed captions still need the same styling and safe-zone work afterwards.

Treat auto captions as a first draft that is ninety percent right. Your job is not to proofread it. Your job is to find the five words that matter, fix them, and change one thing about how you record so those five words behave next time.

Try UGCut

See what UGCut can do and explore every feature: script to posted in about 15 minutes, every cut anchored to a sentence in your script, all on your iPhone.

App overviewComing soon to the App StoreSupport

Frequently Asked Questions

Why do auto captions always misspell my brand name?
Invented names sit outside the recognizer's normal vocabulary, so it substitutes the most likely ordinary word that sounds similar. You cannot make an unknown word known, but you can say it slower, pause before it, and keep it away from music and packaging noise.
Which caption errors are actually worth fixing?
Fix anything that changes meaning, such as a reversed negative or a wrong number, and anything the client will check, such as the brand name or a discount code. Ignore punctuation, capitalisation, and homophones that read fine at speed.
Does a better microphone fix caption accuracy?
It helps by improving the ratio between your voice and the room, but it does nothing for unknown words, homophones, or numbers spoken too quickly. Treat it as one improvement among several rather than a fix.
When should I re-record instead of correcting the text?
When a whole sentence comes back as nonsense, the audio is the problem, not the transcript. If your workflow keeps takes per script line, a two-second retake of that one sentence is faster and cleaner than typing a repair.

Written by

The UGCut Team

Creators & Makers at UGCut

The UGCut team builds the script-first video studio for UGC creators on iPhone. We ship videos ourselves (scripts, hooks, takes, captions, exports), so the blog is the playbook we actually use: how to write a hook that holds, how to record a clean take in one pass, and how to go from script to a platform-ready video in about fifteen minutes without your footage ever leaving your phone. Everything here is written for creators by people who edit on the same phone you do.