Guides

How speech to text works, explained for Mac users

Speech to text works in two moves: the sound of your voice is turned into a picture of its pitches over time, and a trained model reads that picture and predicts the words that match it, one small piece at a time. OpenAI's own description of Whisper follows the same two moves: input audio is split into 30-second chunks, converted into a log-Mel spectrogram and passed into an encoder, and a decoder is trained to predict the text.

I built YapToText, a free Mac dictation app that runs Whisper on your Mac, because my hands make typing difficult and I dictate every day. This page explains the model in plain words, using only what OpenAI and Apple say about their own systems. Then it follows one dictation through YapToText from the microphone to your cursor, and ends with why errors and hallucinations happen.

Download YapToText on the Mac App StoreDownload for free. No account, nothing locked, macOS 14 or later on Apple silicon. Source on GitHub.

On this page: 11 sections
  1. Steps
  2. From sound to a picture of sound
  3. How Whisper turns the picture into words
  4. What Whisper learned from
  5. Where Apple Dictation fits
  6. One dictation through YapToText, end to end
    1. Step 1: The microphone
    2. Step 2: Whisper, on your Mac
    3. Step 3: Commands and dictionaries
    4. Step 4: The cleanup model
    5. Step 5: Insertion
  7. Read each stage in History
    1. Step 6: Open the receipts
  8. Why errors and hallucinations happen
    1. Misheard words
    2. Hallucinations
  9. What it cannot do
  10. The short version
  11. Questions

Steps

  1. Click where the words go, tap Right Command and talk; the audio is cleaned up as it comes in.
  2. Tap Right Command again; Whisper transcribes the audio on your Mac.
  3. Spoken punctuation and your dictionary entries are applied to the raw transcript.
  4. Auto mode or your chosen mode decides whether the cleanup model rewrites the text.
  5. Intelligent Insert fits the text to the words around your cursor and pastes it.
  6. Open History and click Show pipeline details to see each stage.

From sound to a picture of sound

A microphone turns your voice into a long stream of numbers, many thousands every second, each one a measure of the sound at that instant. A speech model does not work on that stream directly. Whisper's paper says all audio is re-sampled to 16,000 Hz, and a log-magnitude Mel spectrogram is computed on 25-millisecond windows with a stride of 10 milliseconds.

In plain words: every 10 milliseconds, Whisper's preparation step takes a short slice of sound and measures how much energy sits at each pitch. Mel is a pitch scale spaced roughly the way human ears tell pitches apart, and log means loudness is measured in ratios, the way ears judge it. Line the slices up and you get a picture, with time running across and pitch running up, in which vowels, hisses and pauses each have their own shape. A speech model does not listen to sound; it reads a picture of it.

Python: The step that makes the picture, from the example in OpenAI's Whisper README.

# load audio and pad/trim it to fit 30 seconds
audio = whisper.load_audio("audio.mp3")
audio = whisper.pad_or_trim(audio)

# make log-Mel spectrogram and move to the same device as the model
mel = whisper.log_mel_spectrogram(audio, n_mels=model.dims.n_mels).to(model.device)

What happens: The recording is cut or padded to 30 seconds and turned into a log-Mel spectrogram, the picture the model reads. You never need to run this to use Whisper in an app; it shows a step that happens inside every run of Whisper, whatever app runs it.

How Whisper turns the picture into words

OpenAI's README calls Whisper a Transformer sequence-to-sequence model. A sequence goes in, the spectrogram, and a different sequence comes out, the text. Its own description, in order:

  1. A 30-second window. The README says its transcribe method processes audio with a sliding 30-second window, so a long recording is read one window at a time.
  2. The spectrogram. Each window becomes a log-Mel spectrogram, as above.
  3. The encoder. OpenAI says the spectrogram is passed into an encoder, the half of the model that works through the sound and hands its result to the other half.
  4. The decoder. It predicts the text, intermixed with special tokens that tell the model which task to do, such as language identification, phrase-level timestamps, multilingual transcription or translation into English.

The README calls the predictions autoregressive: each new piece of text is predicted from the audio plus the text written so far, the way you might finish a sentence you have half heard. A token is one of those pieces, often part of a word. Because all the tasks are written as tokens, the README says a single model can replace many stages of a traditional speech-processing pipeline.

Whisper also punctuates ordinary speech by itself, so commas, periods and capitals arrive with the words. And because every word is a prediction, it can guess. Whisper does not look sounds up in a list; it predicts the most likely text for what it hears. The Whisper for Mac guide covers the model as an app sees it.

What Whisper learned from

A model like this learns from examples, and OpenAI says what Whisper learned from. Its model card says the models are trained on 680,000 hours of audio and the matching transcripts collected from the internet, with the non-English part covering 98 languages. For large-v3, which turbo is built from, OpenAI's Hugging Face card gives 1 million hours of weakly labeled audio and 4 million hours of pseudo-labeled audio collected using large-v2. The README calls Whisper a general-purpose speech recognition model trained on a large dataset of diverse audio, and also a multitasking model: it can recognize speech in many languages, translate speech and identify the language.

  • What the scale buys. The model card says the models show improved robustness to accents, background noise and technical language, compared with many existing systems.
  • What it does not buy. The README says performance varies widely depending on the language, and the model card adds disparate performance on different accents and dialects.
  • Turbo. OpenAI calls turbo an optimized version of large-v3 that offers faster transcription with a minimal degradation in accuracy, and its model card on Hugging Face says the decoding layers were cut from 32 to 4.

YapToText runs Whisper Large v3 Turbo (Q5), the standard 5-bit whisper.cpp file, 574 MB, which I have not fine-tuned or retrained. It runs through a copy of whisper.cpp built with Metal and Accelerate, and whisper.cpp's README calls Apple silicon a first-class citizen. The AI Models page offers ten more Whisper models: bigger ones hear better, and smaller ones are lighter on battery and memory. The model is stock; what changes from app to app is everything around it. The Whisper Large v3 Turbo guide and the AI Models guide have the details.

The AI Models page in YapToText: Whisper Large v3 Turbo (Q5) and Phi-3.5 Mini Instruct marked Included, model library below.
Figure 1. The model that does the hearing in YapToText, Whisper Large v3 Turbo (Q5), under Your defaults, with the Dictation Model Library below it.

Where Apple Dictation fits

Apple Dictation is Apple's own speech recognition, built into macOS. Apple's guides describe what it does and where it runs, not which model it uses inside, so this section sticks to those.

  • Where it runs. Apple's Mac User Guide for macOS Sonoma says that on a Mac with Apple silicon, dictation requests are processed on your device for supported languages, with no internet connection required, and that on an Intel-based Mac, or in a language without on-device support, your dictated words are sent to Apple. Today the text below Dictation in System Settings, Keyboard says whether yours is processed on your Mac or sent to Siri servers.
  • Punctuation. In supported languages it adds commas, periods and question marks on its own, and you can say a mark's name.
  • When it stops. It stops by itself after 30 seconds without speech.
  • A newer model. Apple lists improved accuracy in Dictation from a new on-device model in macOS 27, on Mac models with M3 or later and at least 12 GB of unified memory, marked as available in English.

YapToText does not use Apple Dictation; it runs Whisper. Both turn speech into text, but only one of them lets you choose the model and see what it heard. The Apple Dictation vs Whisper comparison goes engine against engine.

One dictation through YapToText, end to end

Whisper is one step. A dictation in YapToText passes through several more on your Mac, and none of them sends anything over the network. The model was the easy part; the work went into everything around it.

Step 1: The microphone

  • You: Click where the words go, tap Right Command and talk.
  • YapToText: Keeps the half second before you pressed the key, so first words are not clipped. It measures the room's noise and subtracts steady noise only when the signal-to-noise ratio calls for it, lifts a quiet voice relative to the room, and guards the peaks so one loud word cannot clip.
  • Check: The wave in the floating panel moves when you speak. With Live transcription preview on, in Settings, Advanced, Text output, which is the default, your words form under it as you talk.
App Store poster headed It hears you through anything, listing the audio steps above the expanded recording panel.
Figure 2. What YapToText does around Whisper, from background noise removal and peak-guarded capture to deeper decoding in noise.

Step 2: Whisper, on your Mac

  • You: Tap Right Command again.
  • YapToText: Compresses long silences so the decoder does not drift, and hands Whisper a short prompt of about 200 characters, your dictionary entries first, then distinctive words from your recent dictations, which leans it toward your spellings. Decoding switches from greedy, which takes the likeliest next token each time, to beam search, which weighs several candidates, when the clip measures noisy. A rescue pass runs if the transcript comes back shorter than the speech that went in, and a guard collapses repetition loops on silent tails.
  • Check: The wave winds into a spinning ring while your Mac works. If the dictation was only silence, YapToText says "Didn't catch that" and types nothing.

The language comes from the Language picker on the AI Models page: Auto-detect or one of 29 languages, starting on your Mac's system language. A dictation longer than 4 minutes is cut at quiet moments into pieces of about 2 minutes, transcribed as they finish.

Step 3: Commands and dictionaries

  • You: Nothing, unless you said a mark's name, such as question mark at the end of a clause.
  • YapToText: Turns spoken punctuation into marks before any AI cleanup, so the cleanup model sees the finished mark, and applies your dictionary entries to the raw transcript.
  • Check: In History, the raw transcript shows what Whisper heard before any of these fixes ran.

The spoken punctuation guide lists every name, and the Dictionaries guide shows how an entry works twice.

Step 4: The cleanup model

  • You: Nothing, or press a mode's number key while you talk, or end with an instruction.
  • YapToText: With Auto mode on, the default, it reads four signals, cheapest first: a spoken instruction at the end, the shape of the words, the app you are dictating into, and only if it is still unclear, a one-word check by the cleanup model. Every mode but Raw Transcription then hands the transcript to the cleanup model, Phi-3.5 Mini Instruct by default, with that mode's instructions.
  • Check: If the result reads like an assistant's reply, grows well past what you said or repeats the reference text, the app throws it away and types the raw transcript instead.

Say: At the very end of any dictation while Auto mode is on.

make that formal

What happens: The phrase is not typed. Auto mode rewrites the dictation in a formal, professional tone.

After cleanup, your dictionary entries run again, so cleanup cannot undo a fix, and emoji and snippets go in after cleanup, so a rewrite cannot drop or move them. The Auto mode guide has every rule.

Step 5: Insertion

  • You: Nothing.
  • YapToText: Intelligent Insert briefly selects and copies a few words on each side of your cursor with invisible keystrokes, then fits the edges: a lowercase first word mid-sentence, one space where words would collide, no trailing period when the sentence carries on. Then it pastes the text and puts your clipboard back.
  • Check: The words are at your cursor, and the capybara's speech bubble in the menu bar flashes green. Without Accessibility, the words go to the clipboard instead.

The Intelligent Insert guide explains the read and the beep some apps play at it.

Read each stage in History

Every dictation is filed in History, on your Mac, and its pipeline details show each stage above. When a dictation comes out wrong, the pipeline details show which stage to blame.

Step 6: Open the receipts

  • You: Open the History page, find the dictation and click Show pipeline details.
  • YapToText: Shows what the speech model heard, the text after cleanup, dictionaries and commands, and what was actually delivered when that differed.
  • Check: Below them are the timestamp, the mode, what Auto detected, the cleanup model, the delivery outcome, the target app, how long you spoke, the language, and stop to text split into transcription, AI cleanup and delivery.
A YapToText History entry with pipeline details open: raw transcript, cleaned transcript, mode, delivery, target app and timings.
Figure 3. The raw transcript, with its um and lowercase thursday, sits above the cleaned sentence, so you can see what Whisper heard and what the cleanup changed.

By default the recording is kept too, so the play button settles whether you said it or the speech model misheard it. Copy Raw Transcript, in each entry's export menu, gives you exactly what Whisper wrote. The History guide covers the rest.

Why errors and hallucinations happen

Two different things go wrong, and they have different causes. A misheard word has sound behind it, and the model picked the wrong text for it. A hallucination is text with no sound behind it at all.

Misheard words

  • Names and jargon. OpenAI's Whisper prompting guide says it may get uncommon proper nouns wrong, such as names of products, companies or people, and that a prompt with the right spellings can help, though it calls such techniques not especially reliable. That is why YapToText's dictionary entries both prime Whisper and fix what slips through.
  • Language and accent. OpenAI says performance varies widely by language, and its model card notes disparate performance on different accents and dialects.
  • A faint or noisy signal. Less clear sound gives the prediction less to anchor each word to.

Hallucinations

OpenAI's model card explains them. Because the models were trained in a weakly supervised way on large-scale noisy data, their predictions may include text that is not actually spoken; OpenAI's hypothesis is that the model combines trying to predict the next word with trying to transcribe the audio. The model card also says the sequence-to-sequence design makes it prone to repetitive text, which beam search and temperature scheduling mitigate to some degree but not perfectly, and that both may be worse in lower-resource languages. Whisper's paper says its remaining errors, particularly in long-form transcription, include problems such as getting stuck in repeat loops, not transcribing the first or last few words of an audio segment, or complete hallucination. When there is little sound to go on, the predicting half of the model does more of the work. The cleanup model can misbehave too, which is what the check in Step 4 is for. The Whisper hallucinations guide goes deeper.

Climb this ladder, smallest step first:

  1. Read the raw transcript. In History, Show pipeline details. If the wrong text is there, Whisper wrote it; if not, the cleanup did.
  2. Teach it the word. Put the spelling Whisper produced in Heard and yours in Replace with on the Dictionaries page.
  3. Set the language. Pick the language you speak in the Language picker instead of leaving Whisper to work it out.
  4. Give it less silence. Tap the key when you finish, or turn on Auto-stop after silence, off by default, which ends a dictation after 0.5 to 5 seconds of quiet.
  5. Fix the microphone. Pick the Mac's own microphone over a Bluetooth one in the Input picker, and move closer.
  6. Try a bigger model. Whisper Large v3 Turbo (Q8) is a bigger download, Whisper Large v3 is the largest and slowest, and the stars on the AI Models page show the trade.

The guide to improving dictation accuracy works through each fix.

What it cannot do

  • Accuracy numbers. I do not publish accuracy percentages for any model, and nothing here measures one engine against another.
  • Every hallucination. The guards around Whisper are measured against a harness of my own real dictations before they ship. That is one person's recordings, not a benchmark, and a new kind of mistake can still get through.
  • Translation as you speak. It types what you said in the language you said it, and OpenAI notes the turbo model is not trained for translation tasks.
  • Speakers and timestamps. File transcription gives plain text, with no timestamps, subtitles or speaker labels.
  • Intel Macs. YapToText needs Apple silicon and macOS 14 or later, and nothing is claimed for Intel Macs.

Knowing how it works does not make it perfect; it tells you where to look when it is not.

The short version

  • The picture. Audio becomes a log-Mel spectrogram, a map of pitch over time.
  • The prediction. Whisper's encoder reads it and its decoder predicts the text, one token at a time, 30 seconds at a time.
  • The pipeline. In YapToText, cleaner audio goes in, and commands, dictionaries, an optional cleanup model and Intelligent Insert shape what comes out.
  • The mistakes. Prediction is why names slip and why silence can produce words; the raw transcript shows which.

YapToText is free on the Mac App Store, and every step above runs on your Mac.

Speech to text is a very good guess at what you said, so read what lands. If any step here does not match what you see in History, open an issue on GitHub. Every one gets read.

Download YapToText on the Mac App StoreDownload for free. No account, nothing locked, macOS 14 or later on Apple silicon. Source on GitHub.

About the author. Ryleigh Newman is a developer whose hands make typing difficult and who dictates every day with YapToText, the app this site describes. Facts about other apps come from their own sites, with the date they were read. More about the project.

Sources

Questions

How does speech to text work?

The audio is turned into a spectrogram, a picture of pitch over time, and a trained model predicts the matching text one piece at a time. OpenAI's Whisper splits audio into 30-second chunks, converts them to a log-Mel spectrogram, and uses an encoder and a decoder to produce the text.

How does Whisper work?

OpenAI describes Whisper as a Transformer sequence-to-sequence model trained on 680,000 hours of audio and transcripts. Its encoder reads the log-Mel spectrogram of each 30-second window, and its decoder predicts the text, with special tokens for the language and the task.

How does dictation work on a Mac?

Apple Dictation turns what you say into text at your cursor. Apple's Sonoma guide says it is processed on the device for supported languages on Apple silicon, and the text below Dictation in Keyboard settings says whether yours goes to Siri servers. Apps such as YapToText run their own model, Whisper, on the Mac instead.

Why does speech to text get words wrong?

It predicts the most likely text for the sound, so unusual names, noise, faint speech and less common languages and accents give it room to guess wrong. Whisper can also produce text nobody said, which OpenAI's model card calls hallucination.

Does speech to text need the internet?

Not when the model runs on the Mac. YapToText's App Store build ships Whisper and the cleanup model, so it works offline. Apple Dictation may need a connection, depending on your Mac and language.

What does the cleanup model do after Whisper?

In YapToText, every mode but Raw Transcription hands Whisper's text to Phi-3.5 Mini Instruct on your Mac, which shapes it for that mode. If the result looks like an assistant's reply or grows well past what you said, the raw transcript is typed instead.