Skip to content
WitnessCompare apps

Dictation that drops words

Why dictation silently drops words, and how to catch it

Short answer

The short answer: dictation does not fail loudly. Speech recognition predicts language as much as it transcribes it, so when the audio is unclear it does not leave a gap, it writes the most probable words instead. OpenAI states this plainly in the Whisper model card: the predictions may include texts that are not actually spoken in the audio input. A peer-reviewed study of Whisper transcriptions found that roughly 1% contained entire hallucinated phrases or sentences which did not exist in any form in the underlying audio. The failure you need to catch is not the garbled sentence, because you already notice that one. It is the clean sentence where a figure, a date, a name or the word not came out different and the text still reads perfectly.[2][3]

Try free on your Mac

For Mac users who write a lot every day and mix English with another language. Runs on Apple silicon.

Facts checked on . Prices in their original currency; no accuracy benchmarks of our own.
01 / Overview

How dictation fails, and why the quiet failures are the expensive ones

What goes wrongWhat you see on screenWhat it can cost
A negation disappearsA fluent sentence with the opposite meaningAn approval where you dictated a refusal
A figure comes out differentA plausible number in the right placeThe wrong amount in a quote or an invoice
A date shiftsA valid date, just not the one you saidA deadline nobody notices until it passes
A name is normalised to a common oneA name that looks rightThe wrong person in a record
The model writes a phrase you never saidText with no basis in the audio at allA statement attributed to someone who never made it
Text repeats in a loopAn obviously broken paragraphLittle, because this is the one you do catch
02 / Details

Recognition guesses, and a guess reads like a fact

A speech model does not hand back silence when it cannot make out a word. It is trained to produce the most likely sequence of words, so an unclear passage comes out as fluent text rather than as a gap or a question mark. OpenAI documents this in the Whisper model card under limitations: the predictions may include texts that are not actually spoken in the audio input, which the card calls hallucination, and the sequence-to-sequence architecture also makes the model prone to generating repetitive texts.[2]

That is why the dangerous error is the tidy one. A mangled sentence announces itself. A sentence that reads perfectly but says 14,000 instead of 40,000, or drops a single not, gives you nothing to notice. You reread your own dictation, it matches what you meant to say, and you send it.[2]

03 / Details

How often this happens, and to whom

The best public measurement comes from Careless Whisper, presented at the ACM conference on Fairness, Accountability and Transparency in 2024. Across the transcriptions the authors studied, roughly 1% contained entire hallucinated phrases or sentences which did not exist in any form in the underlying audio, and 38% of those hallucinations included explicit harms such as perpetuating violence, making up inaccurate associations or implying false authority.[3]

One percent sounds survivable until you count what a working day contains. It also is not spread evenly. The same study found hallucinations occurred disproportionately for speakers with longer stretches of non-vocal time, a common symptom of aphasia, which means the people least able to proofread the output are the ones getting the most invented text. OpenAI's own card reports uneven performance across languages, accents and dialects for the same underlying reason.[3][2]

04 / Details

Running it locally does not make it immune

This is worth saying clearly, because local dictation is often sold as the safe option and we sell local dictation. Keeping recognition on your own machine changes where the audio goes. It does not change how the model behaves when the audio is unclear. An offline model drops a negation exactly the same way a cloud model does, and a private mistake is still a mistake.[2]

So the honest question is not which app recognises speech most accurately. There is no shared published benchmark with identical hardware, recordings and post-processing that would let anyone rank dictation apps on accuracy, and we do not claim to win such a comparison. The question is what the app shows you when it was unsure.[1]

05 / Details

Automatic cleanup removes the evidence

Most dictation tools now add a polishing pass: grammar fixed, filler words removed, punctuation inserted, the whole thing smoothed into prose. For most writing that is exactly what you want. It also deletes the tells. Hesitation, a half-repeated word, an odd clause, these are the things that make you stop and reread, and a cleanup pass is designed to remove precisely them.[1]

That is the trade nobody mentions: the smoother the output, the less chance you have of spotting that the amount changed on the way in. If a tool rewrites your sentence, you want to be able to see the untouched version next to it.[1]

06 / Details

What you can do today, with no new software

macOS already helps more than people realise. Ambiguous text is underlined in blue, and clicking the underlined word offers alternatives, so those marks are worth looking for rather than scrolling past. Beyond that, three habits catch most of it: read every number and date out of the screen rather than from memory, search the text for not, no and never before sending anything consequential, and dictate names and figures in short separate passes instead of inside long sentences.[4]

None of that scales to a full working day, which is the real problem. Proofreading your own dictation word by word costs more time than the dictation saved, so in practice people stop doing it after the first week and the errors go out unseen.[1]

07 / Details

What catching it reliably requires

Four things, and they are specific. Mark the categories where an error changes meaning rather than flagging general uncertainty: figures, dates, proper names, negations and your own terminology. Keep the unmodified raw transcript available, so you can see what the cleanup changed. Let the person hear the short piece of audio behind a marked spot, because reading the text again cannot tell you what you actually said. And narrow the set of languages the model may choose from, so a word is not rendered into a language you were not speaking.[1]

Witness is built around those four. The risky categories are marked before the text is inserted, a hotkey shows the raw transcript, a marked fragment can be played back from memory, and you tick which of 100 languages you actually speak instead of letting the model guess per segment. German, English, Russian and Ukrainian are measured end to end, from recognition through cleanup to every risk marking; for the other languages, markings that need per-language calibration stay off rather than guessing. Everything runs on an Apple Silicon Mac and the audio is not written to disk. The trial is 13 days of the full version with no account.[1]

Which one fits

What we recommend, and where the line is

What we recommend: treat the clean sentence as the one to check, not the garbled one. Turn on the built-in dictation, learn to look for the blue underlines, and read figures and dates off the screen before anything consequential leaves your hands. If you dictate all day into documents other people act on, that discipline will not hold, and then the thing worth paying for is an app that marks the risky spots for you and can still show you the raw transcript and the original audio. That is what Witness does, and it is the only reason to choose it over the free local alternatives.[4][1]

Try free on your Mac
FAQ / Short answers

Frequently asked questions

Why does dictation leave out words like not?

Because the model predicts the most probable sequence of words rather than reporting that it was unsure. A short unstressed word carries little acoustic signal, so it is the first thing to go, and the sentence left behind is still grammatical. That is what makes it hard to notice.

What is a speech-to-text hallucination?

Text in the transcript that was never spoken. OpenAI's Whisper model card states that predictions may include texts that are not actually spoken in the audio input. A 2024 study at ACM FAccT found roughly 1% of the transcriptions it examined contained entire hallucinated phrases or sentences absent from the audio.

How often does Whisper hallucinate?

In the Careless Whisper study, roughly 1% of transcriptions contained entirely hallucinated phrases, and 38% of those hallucinations included explicit harms. Rates were higher for speakers with longer non-vocal stretches, a common symptom of aphasia.

Does offline or local dictation avoid the problem?

No. Local processing changes where your audio goes, not how the model behaves when the audio is unclear. An offline model drops a negation the same way a cloud model does.

How can I tell whether dictation changed a number?

Read it off the screen rather than from memory, and keep a way back to the original. A tool that shows the unmodified raw transcript and can replay the audio behind a particular spot answers the question directly; without one, the only method is proofreading every figure.

Is any dictation app more accurate than the others?

Nobody can answer that from public evidence. There is no shared published benchmark with identical hardware, recordings and post-processing covering these apps, so any accuracy ranking, including one in our favour, would be unfounded.

Sources

Official sources

Only official product, help, pricing and legal pages. Every source was retrieved on 1 October 2026.

  1. WitnessWitness product page
  2. OpenAIWhisper model card, Limitations and biases
  3. Koenecke, Choi, Mei, Schellmann, SloaneCareless Whisper: Speech-to-Text Hallucination Harms (ACM FAccT 2024)
  4. Apple SupportDictate messages and documents on Mac