Whisper hallucinations: why Whisper types words nobody said, and how to stop it
By Ryleigh Newman ยท Published
Whisper hallucinations are words in a transcript that nobody said: a phrase repeated over and over, a sign-off such as "Thank you for watching" over silence, or a whole sentence made up. They happen because Whisper writes text much as a language model does, and when the audio is silent, noisy or unclear, its sense of what usually comes next can outweigh what it actually heard.
OpenAI's own model card warns about it, and a 2024 study found invented phrases in roughly 1% of the transcriptions it tested. I built YapToText, a free Mac dictation app that runs Whisper on your Mac, because my hands make typing difficult and I dictate every day, and a good share of the work around the model went into catching exactly this. Below: what hallucinations look like, why they happen, what the research found, what YapToText does about them, and what you can do with any Whisper setup, whisper.cpp included.
Download for free. No account, nothing locked, macOS 14 or later on Apple silicon. Source on GitHub.
On this page: 11 sections
Steps
- Stop the recording as soon as you finish talking, or turn on Auto-stop after silence on the Dictation page.
- Press Space to pause during a long think instead of recording the silence.
- Choose the language you speak in the Language picker on the AI Models page.
- Before transcribing a file, trim long silences and music, and convert it to 16-bit WAV if you use whisper-cli.
- In History, click Show pipeline details and read the raw transcript, or press 2 while you talk for Raw Transcription.
What a Whisper hallucination looks like
OpenAI's Whisper paper lists the failure modes in plain words: getting stuck in repeat loops, not transcribing the first or last few words of an audio segment, and complete hallucination, where the model outputs a transcript entirely unrelated to the actual audio. In practice they show up as four patterns. A hallucination is not a misheard word; it is a word with no sound behind it at all.
- Repetition loops. The same sentence, phrase or word again and again, often at the end of a recording that trails off into silence. OpenAI's model card says the sequence-to-sequence architecture makes the model prone to generating repetitive text, which beam search and temperature scheduling can mitigate to some degree, but not perfectly.
- Text from silence. In the Whisper repository's own discussion thread titled "Share your hallucinations here", users report lines such as "Thank you for watching" and subtitle credits appearing over silence, music or sound effects, where nobody was speaking.
- Invented sentences. Whole phrases that were never said. Cornell's write-up of the Careless Whisper study says Whisper occasionally makes up entire phrases and sentences, sometimes with violent language, invented personal information and fake websites.
- Labels and stage directions. Text that reads like a script rather than speech. YapToText's 1.4 release notes record a fix for hallucinated speaker labels such as "Male speaker:" appearing in transcripts.
The opposite failure, dropped words at the start or end, is related, and it is why a missing first word and an invented last line can come from the same recording. Ordinary misheard words are a different problem with different fixes, covered in the Whisper for Mac guide.
Why Whisper hallucinates
OpenAI explains the cause in its model card. Because the models were trained in a weakly supervised way on large-scale noisy data, their predictions may include text that is not actually spoken in the audio. The model card suggests this happens because the model combines trying to predict the next word with trying to transcribe the audio itself, and it says the behavior may be worse in lower-resource or lower-discoverability languages. When there is little sound to go on, the guessing half of the model does more of the work.
That points to the conditions that invite it:
- Silence and long pauses. The Careless Whisper study found hallucinations occur disproportionately for people who speak with longer shares of non-vocal time, longer pauses in other words.
- Music and sound effects. The reports in the Whisper repository cluster around stretches with no speech in them.
- A noisy or faint signal. Less clear speech gives the model less to anchor each word to.
- The language. OpenAI says performance varies widely depending on the language, and the model card says hallucination may be worse in lower-resource languages.
One more detail matters for spotting it. The study found that hallucinations from OpenAI's Whisper API were non-deterministic, yielding different random text on each run, and used differences between two runs of the same audio to find them. If your tool gives different text for the same audio on two runs, the part that changed deserves a second look.
What the Careless Whisper study found
The study is "Careless Whisper: Speech-to-Text Hallucination Harms", by Allison Koenecke, Anna Seo Gyeong Choi, Katelyn X. Mei, Hilke Schellmann and Mona Sloane, presented at the ACM Conference on Fairness, Accountability, and Transparency, FAccT 2024. The rate is small; the harm in a single invented sentence can be large.
Checked October 5, 2026, in the paper on arXiv and the ACM Digital Library, and in Cornell's news story.
- What they tested. Nearly 40 hours of audio from AphasiaBank, 13,140 segments, from speakers with aphasia and a control group, sent to OpenAI's Whisper API in 2023.
- How often. Roughly 1% of transcriptions contained entire hallucinated phrases or sentences that did not exist in any form in the underlying audio.
- How bad. 38% of the hallucinations included explicit harms, such as perpetuating violence, making up inaccurate associations, or implying false authority.
- Who. Speakers with aphasia got more hallucinations than the control group, and segments with more non-vocal time were more likely to produce them.
- Over time. When the authors reran the audio in December 2023, many examples were resolved, but 12 of 187 audio segments still produced hallucinations.
Two things to keep straight when you read those numbers. They describe OpenAI's hosted API on that sample in 2023, not whisper.cpp, not Whisper Large v3 Turbo and not YapToText, so they are not a forecast for your own dictation. And they are not a reason to panic: they are a reason to read what you dictated, which is good practice anyway. If long pauses are part of how you speak, dictation by condition says honestly where YapToText may fall short.
How YapToText guards against it
YapToText runs the stock Whisper Large v3 Turbo (Q5) file through whisper.cpp; I have not retrained the model. What I did build is a pipeline around it, and part of that pipeline exists to keep words nobody said out of your document. The model is stock; the guards around it are where the work went.
Before Whisper hears anything
- Long silences are compressed before transcribing, so the model does not drift.
- Noise is measured on every dictation, and steady noise is subtracted only when the ratio calls for it. A quiet voice is lifted relative to the room.
- Noisy clips get a deeper search. Decoding switches from greedy to beam search when the clip measures noisy.
- The half second before you press the key is kept, so first words are not clipped.

After the transcript comes back
- Repetition loops are caught and removed. A guard collapses the loops Whisper falls into on silent tails, and since 1.4 a long dictation that ends in silence no longer repeats its last sentence over and over.
- Invented speaker labels are removed, along with stage directions. Before the 1.4 fix, a label like "Male speaker:" could even be learned as vocabulary.
- Pleasantries hallucinated from silence are caught and removed.
- Silence types nothing. If a dictation was only silence, YapToText says "Didn't catch that" and nothing is transcribed.
- A short transcript gets a second pass. If the text comes back shorter than the speech that went in, a rescue pass runs.
When the cleanup model misbehaves
Every mode except Raw Transcription hands the transcript to a small cleanup model on your Mac, and a language model can invent too. The app checks its work: if the result reads like an assistant's reply, grows well past what you said, repeats the reference text or invents an ellipsis, it is thrown away and the raw transcript is typed instead. The cleanup is also told never to drop a sentence or add anything you did not say. The AI dictation page covers those rules.
Every one of these changes is measured against a harness of my own real dictations before it ships. That is one person's recordings, not a benchmark, and caught and removed is not the same as never: a new kind of hallucination can still slip through.
See what Whisper actually heard
The quickest way to tell a hallucination from a cleanup mistake is to look at what Whisper produced before anything else touched it. If the invented text is in the raw transcript, Whisper wrote it; if not, the cleanup did.
Step 1: Open the dictation in History
- You: Open History, find the dictation and click Show pipeline details.
- YapToText: Shows the raw transcript above the final text, with the mode, the language and the timings below.
- Check: Read the raw transcript against what you remember saying.

Step 2: Dictate the next one raw
- You: While you talk, press 2, Raw Transcription in the default order with Auto mode on.
- YapToText: Types Whisper's words with basic punctuation and no AI model, for that one dictation.
- Check: What lands is what Whisper heard. To keep every dictation raw, turn off Post-transcription analysis on the Home page.
Copy Raw Transcript, in each History entry's export menu, puts the version from before cleanup on the clipboard. And if you want to read every long dictation before it lands, turn on Review before inserting in Settings, Advanced, Text output. The modes guide explains what each mode is allowed to change.
Prevent hallucinations in any Whisper app
Most hallucinations grow out of silence, so the best prevention is giving Whisper less of it. These work with any Whisper app, and the YapToText settings are named where they apply. Less silence in, fewer invented words out. Climb this ladder, smallest step first:
- Stop when you are done. Tap the key as soon as you finish talking, instead of leaving the recording running. In YapToText, Auto-stop after silence, off by default in the Recording card on the Dictation page, ends a dictation after 0.5 to 5 seconds of quiet.
- Pause on purpose. For a long think in the middle, press Space to pause the dictation and again to resume. The space is never typed.
- Set the language. Choose the language you speak in the Language picker on the AI Models page instead of leaving Whisper to work it out. The languages page lists all 29.
- Give it a clean signal. Microphone close, room quiet, music off. YapToText's Pause music and video while dictating, on by default, pauses playback while you talk.
- Trim recordings first. Before you transcribe a file, cut long silent stretches and music from the start and end in any audio editor.
- Read before it matters. Read the raw transcript for anything that will be sent, filed or quoted.

The keys and shortcuts page covers Space, Esc and the rest.
Tips for whisper.cpp users
If you run Whisper yourself through whisper.cpp, you control the audio that goes in, which is where most of the prevention lives. Clean, trimmed audio is the cheapest hallucination fix there is.
- Convert properly. whisper.cpp's README says whisper-cli currently runs only with 16-bit WAV files, and gives an ffmpeg command for converting anything else.
- Trim first. Cut long silences and music before you convert, so the model never sees them.
- Set the language if you can. Check whisper-cli's help output for a language option rather than relying on detection.
- Try another model size. If you only speak English and memory is tight, OpenAI says the English-only .en models tend to do better for English, most of all at tiny and base.
- Know what live mode is. whisper-stream samples the microphone every half a second and transcribes continuously, and the README calls it a naive example.
Terminal: Convert a trimmed recording to the WAV format whisper-cli reads, with the README's own command.
ffmpeg -i input.mp3 -ar 16000 -ac 1 -c:a pcm_s16le output.wavWhat happens: ffmpeg writes output.wav as 16 kHz, mono, 16-bit audio, which you then hand to whisper-cli after -f.
A model file you downloaded for whisper.cpp also works in YapToText: under Your own models on the AI Models page, Add Model takes any whisper.cpp GGML .bin file. The local Whisper guide has the full setup, model sizes and memory figures.
What no guard can promise
- Every hallucination. The guards catch the patterns I have seen and tested for. A new one can get through, which is why the raw transcript is always one click away.
- Words never heard. No guard can recover speech the microphone did not catch, and removing invented text leaves a gap, not the real words.
- A rate of my own. I have not measured a hallucination rate for YapToText or for any model, so this page quotes only the published study, with its limits.
- File transcription extras. Files come back as plain text, with no timestamps, so there is no jump from a suspicious line to its place in the audio. The audio file page covers that feature.
- Long pauses as a speaking style. Whisper is a general-purpose model, and nothing in YapToText changes how it handles speech with long pauses beyond compressing the silence.
- An Intel Mac. YapToText needs an Apple silicon Mac with macOS 14 or later.
Guards lower the odds; reading what you dictated is still the last check.
The short version
- What it is. Text with no speech behind it: loops, sign-offs over silence, invented sentences and labels.
- Why. Whisper predicts as well as transcribes, and silence, pauses, music and noise give the prediction room.
- In YapToText. Long silences compressed; repetition loops, speaker labels and pleasantries from silence removed; an overgrown cleanup thrown away.
- For anyone. Stop when you finish, trim silence, set the language, and read the raw transcript.
YapToText is free on the Mac App Store, and Whisper runs on your Mac.
Give Whisper less silence, and check what it wrote. If YapToText ever types something you did not say, open an issue on GitHub with the raw transcript from History. Every one gets read.
Download for free. No account, nothing locked, macOS 14 or later on Apple silicon. Source on GitHub.
Related
Sources
- https://arxiv.org/abs/2402.08021 (checked 2026-10-05)
- https://dl.acm.org/doi/10.1145/3630106.3658996 (checked 2026-10-05)
- https://news.cornell.edu/stories/2024/06/ai-speech-text-can-hallucinate-violent-language (checked 2026-10-05)
- https://github.com/openai/whisper/blob/main/model-card.md (checked 2026-10-05)
- https://cdn.openai.com/papers/whisper.pdf (checked 2026-10-05)
- https://github.com/openai/whisper/discussions/1873 (checked 2026-10-05)
- https://github.com/openai/whisper (checked 2026-10-03)
- https://github.com/ggml-org/whisper.cpp (checked 2026-10-03)
- https://github.com/ryleighnewman/YapToText (checked 2026-10-03)
- https://apps.apple.com/us/app/yaptotext/id6786382289 (checked 2026-10-03)
Questions
Why does Whisper hallucinate?
OpenAI's model card says Whisper was trained on large-scale noisy data and may output text that was not spoken, likely because it predicts the next word as well as transcribing. Silence, long pauses, music and noise give that prediction more room.
Why does Whisper write "thank you for watching"?
It is a known kind of hallucination over silence or music; users in the Whisper repository's "Share your hallucinations here" thread report it along with subtitle credits. Stopping the recording when you finish talking and trimming silence give it less room.
Why does Whisper keep repeating the same text?
OpenAI's paper lists getting stuck in repeat loops as a failure mode, and its model card says the architecture is prone to repetitive text. Loops often appear on a silent tail, which YapToText's guard collapses.
How common are Whisper hallucinations?
The Careless Whisper study found roughly 1% of transcriptions from OpenAI's Whisper API in its 2023 sample contained invented phrases or sentences. Your own rate depends on your audio, your model and your app.
How do I fix Whisper hallucinations?
Give it less silence: stop when you are done, pause on purpose and trim recordings. Set the language, use a clean microphone signal, and read the raw transcript so you can see what Whisper actually wrote.
Does YapToText hallucinate?
It runs Whisper, so it can. It compresses long silences, removes repetition loops, invented speaker labels and pleasantries hallucinated from silence, and throws away a cleanup that grows well past what you said. History shows the raw transcript.