Local Whisper on a Mac: what it needs, what it costs, and how to run it
By Ryleigh Newman ยท Published
Local Whisper means running OpenAI's Whisper speech model on your own Mac, so your audio becomes text on the machine instead of on a server somewhere. On a Mac, the engine to know is whisper.cpp, which treats Apple silicon as a first-class citizen, and the decision that matters is the model file: bigger models hear better and take more disk space and memory.
This page covers how the pieces fit, what each model size costs by whisper.cpp's own figures, the commands to run it yourself, and where YapToText fits. I built YapToText, a free app that runs whisper.cpp with the stock Whisper Large v3 Turbo (Q5) file, so weigh that part with that in mind. If you only want to know whether there is a Whisper app for Mac, the Whisper Mac app page answers that.
Download for free. No account, nothing locked, macOS 14 or later on Apple silicon. Source on GitHub.
On this page: 9 sections
Steps
- Clone the whisper.cpp repository from GitHub and open its folder in Terminal.
- Download a model with the project's download-ggml-model.sh script, such as large-v3-turbo-q5_0.
- Build the project with CMake, using the two cmake commands from its README.
- Convert your recording to a 16-bit WAV file, for example with ffmpeg.
- Run whisper-cli with your model file and your WAV file, and read the transcript in Terminal.
Know the three pieces of local Whisper
Local Whisper has three pieces, and all three sit on your Mac. The model is a file, the engine runs it, and your audio never has to leave the machine.
- The model. Whisper is OpenAI's speech recognition model, and OpenAI released its code and model weights under the MIT License. You download a model file once, and that file does the listening.
- The engine. whisper.cpp is a plain C and C++ implementation of Whisper, also under the MIT License. Its README calls Apple silicon a first-class citizen, optimized through ARM NEON, Accelerate, Metal and Core ML, and says that on Apple silicon the inference runs fully on the GPU through Metal. OpenAI's own package runs Whisper from Python instead.
- Your audio. A recording or your microphone goes in, and text comes out. The internet is needed only to download the model file in the first place.
That is the whole difference from a cloud service: nothing is uploaded to be transcribed. It also means your Mac does the work, so the model you pick decides how much memory and disk space that work takes.

Pick a model by disk space and memory
whisper.cpp publishes what each model takes on disk and in memory, and OpenAI publishes how many parameters each one has. Turbo is OpenAI's faster version of large-v3, and its 5-bit file takes about a fifth of large's disk space.
Checked October 3, 2026, in the whisper.cpp README and model list and in OpenAI's Whisper README.
| Model | Parameters (OpenAI) | File on disk (whisper.cpp) | Memory (whisper.cpp) |
|---|---|---|---|
| tiny | 39 M | 75 MiB | About 273 MB |
| base | 74 M | 142 MiB | About 388 MB |
| small | 244 M | 466 MiB | About 852 MB |
| medium | 769 M | 1.5 GiB | About 2.1 GB |
| large (v1, v2 or v3) | 1550 M | 2.9 GiB | About 3.9 GB |
| large-v3, 5-bit (q5_0) | Quantized large-v3 | 1.1 GiB | Not listed |
| large-v3-turbo | 809 M | 1.5 GiB | Not listed |
| large-v3-turbo, 5-bit (q5_0) | Quantized turbo | 547 MiB | Not listed |
MiB and GiB count in steps of 1,024, so the 547 MiB turbo file in whisper.cpp's list is the same file YapToText shows as 574 MB. That is how two honest pages can list one file as 547 and 574 and both be right. whisper.cpp's memory table covers the five original sizes and gives no figure for turbo or for the quantized files.

A few more notes from the two projects:
- English-only models. tiny, base, small and medium also come as .en files of the same size that hear English only. OpenAI says they tend to do better for English, especially tiny.en and base.en, and that the difference shrinks for small.en and medium.en.
- Quantized models. Files ending in q5_0 are quantized. whisper.cpp says quantized models need less memory and disk space and, depending on the hardware, can be processed more efficiently.
- Turbo. OpenAI calls it an optimized version of large-v3 with faster transcription and a minimal degradation in accuracy, and rates it at about 8 times the speed of large. That ratio was measured on an A100, not on a Mac, and OpenAI says turbo is not trained for translation.
- Languages. OpenAI says Whisper's performance varies widely depending on the language.
- OpenAI's memory column. OpenAI's README lists required VRAM for its Python package instead: about 1 GB for tiny and base, about 6 GB for turbo and about 10 GB for large. Those are not whisper.cpp numbers.
I chose the 5-bit turbo file for YapToText because, on my own dictations, it heard almost exactly what the full model heard, with a third of the memory. That is one person's dictations, not a benchmark, so try two sizes on your own voice before you settle.
Run Whisper yourself with whisper.cpp
Running whisper.cpp yourself takes Terminal and the two tools its README's commands assume, Git and CMake. These are the README's own commands, with the model swapped for the 5-bit turbo file. What you get is a transcript of an audio file in Terminal, not words typed into your apps.
Step 1: Get the code and a model
Terminal: Clone whisper.cpp and download the 5-bit turbo model.
git clone https://github.com/ggml-org/whisper.cpp.git
cd whisper.cpp
sh ./models/download-ggml-model.sh large-v3-turbo-q5_0What happens: The project lands in a whisper.cpp folder, and the script downloads ggml-large-v3-turbo-q5_0.bin from Hugging Face into its models folder. Swap in base.en or small for a smaller download.
Step 2: Build it and transcribe the sample
Terminal: Build whisper-cli, then run it on the sample recording that comes with the project.
cmake -B build
cmake --build build -j --config Release
./build/bin/whisper-cli -m models/ggml-large-v3-turbo-q5_0.bin -f samples/jfk.wavWhat happens: CMake builds the project, and whisper-cli transcribes the sample with the model you downloaded and prints the text in Terminal.
Step 3: Convert your own audio first
The README says whisper-cli currently runs only with 16-bit WAV files, and it gives an ffmpeg command for converting anything else. ffmpeg is a separate install.
Terminal: Turn an MP3 into the WAV format whisper-cli reads.
ffmpeg -i input.mp3 -ar 16000 -ac 1 -c:a pcm_s16le output.wavWhat happens: ffmpeg writes output.wav as 16 kHz, mono, 16-bit audio, which you then hand to whisper-cli after -f.
Or use OpenAI's Python package
Terminal: Install OpenAI's own package, if you would rather work in Python.
pip install -U openai-whisperWhat happens: pip installs the whisper command. OpenAI's README says the code is expected to work with Python 3.8 to 3.11, it needs ffmpeg, and its default model is turbo.
For live speech, whisper.cpp includes whisper-stream, which samples your microphone every half a second and transcribes continuously. It needs SDL2, and the README calls it a naive example. Getting the words into Mail, Slack or your editor is the part an app adds.
Use the same model in YapToText without Terminal
YapToText runs the same engine and the same file with no Terminal window: a copy of whisper.cpp built into the app with Metal and Accelerate, and ggml-large-v3-turbo-q5_0.bin inside the App Store download. Its SHA-256 matches the file on Hugging Face, so it is the stock model, not a custom-tuned or retrained one. I built the app because my hands make typing difficult, and I wanted Whisper typing into every app instead of printing into a Terminal window. It is the file the script downloads; what changes is everything around it.

Here is what the app adds around the model:
- Memory it gives back. On the Energy page, Keep models in memory decides how long a model stays loaded after a dictation: 2 minutes after use by default, with choices from Only while dictating to Until quit. Unloading sooner saves energy and memory, and the next dictation after an unload takes a few extra seconds.
- Room for macOS. When macOS runs short of memory, YapToText drops its loaded models at once and waits five minutes before warming them again.
- A lighter model on battery. Switch models with the power source, off by default, gives the dictation model and the cleanup model each a plugged-in and an on-battery choice.
- Your own file. Under Your own models on the AI Models page, Add Model takes any whisper.cpp GGML .bin, so a model you downloaded with the script works too. The app reads the file to tell what it is and rejects anything else.
- Your Mac's numbers. The This Mac card on the Energy page shows your chip and memory, which is the number to hold up against the table above.
The Dictation Model Library on the AI Models page has ten more Whisper models, from Tiny to Large v3, each with the size the app shows. The AI models guide lists them, and the energy guide explains the memory settings. The App Store download is about 2.9 GB because both of the app's models are inside it. The Homebrew build is about 12 MB and needs a speech model download before its first dictation.
If local Whisper is slow or mishears you
Start with the cheap fixes before you download a bigger model. Fix the microphone and the words before you change the model. Climb this ladder, smallest step first:
- Wait out the first run. A model has to load before it can transcribe. In YapToText the first dictation after launch takes a moment. If you built whisper.cpp with Core ML, its README says the first run on a device is slow while the Neural Engine service compiles the model.
- Set the language you speak. YapToText's Language picker on the AI Models page starts on your Mac's system language, and a language Whisper cannot use falls back to auto-detect. The languages page lists all 29.
- Fix the microphone. Bluetooth microphones are phone quality, so the Mac's own microphone hears better. In YapToText, pick it in the menu bar panel's Input picker.
- Teach it your words. YapToText's Dictionaries page primes Whisper with your names and terms, not just fixes what it types. The dictionaries guide shows how.
- Use an English-only model for English. If memory is tight and you only speak English, OpenAI says the .en versions tend to do better, most of all at tiny and base.
- Change the model size. Go bigger if it mishears and your memory allows it, or smaller if it is slow. In YapToText, the star ratings on the AI Models page show the trade, and every dictation in History shows its own stop-to-text time.

What local Whisper does not change
Running Whisper locally moves the work onto your Mac, and that is all it moves. Local changes where your voice goes, not how well it is heard.
- Mistakes. The model decides how well it hears, wherever it runs, so a name it mishears needs a dictionary entry or a better microphone, not a different location.
- Translation. OpenAI says turbo is not trained for translation, and YapToText only ever asks Whisper to transcribe, never to translate.
- One download. A local model still has to arrive once. whisper.cpp's script fetches it from Hugging Face, YapToText's App Store build has it inside, and the Homebrew build fetches it from the AI Models page.
- Intel Macs. whisper.cpp lists Intel and Arm Macs as supported. YapToText needs Apple silicon and macOS 14 or later, and nothing is claimed for Intel Macs.
- Accuracy numbers. I do not publish accuracy percentages for any model. The stars and your own History are the guide.
To prove a setup is really local, the offline dictation page has a one-minute test: turn Wi-Fi off and dictate.
The short version
- Local Whisper. A model file and an engine on your Mac, with nothing uploaded.
- The model. The 5-bit turbo file, 547 MiB in whisper.cpp's list, is the one YapToText ships.
- Do it yourself. Clone whisper.cpp, download a model, build with CMake, and run whisper-cli on a 16-bit WAV file.
- Or skip Terminal. YapToText runs the same file and types into any app, for free.
Pick the model by your Mac's memory, then by your own voice. If something here does not match what you see, open an issue on GitHub. Every one gets read.
Download for free. No account, nothing locked, macOS 14 or later on Apple silicon. Source on GitHub.
Related
Sources
- https://github.com/ggml-org/whisper.cpp (checked 2026-10-03)
- https://github.com/ggml-org/whisper.cpp/blob/master/models/README.md (checked 2026-10-03)
- https://github.com/openai/whisper (checked 2026-10-03)
- https://huggingface.co/ggerganov/whisper.cpp/blob/main/ggml-large-v3-turbo-q5_0.bin (checked 2026-10-03)
- https://github.com/ryleighnewman/YapToText (checked 2026-10-03)
- https://apps.apple.com/us/app/yaptotext/id6786382289 (checked 2026-10-03)
Questions
How do I run Whisper locally on a Mac?
Use whisper.cpp: clone the repository, download a model with its download-ggml-model.sh script, build it with CMake, and run whisper-cli on a 16-bit WAV file. Or install an app that bundles it, such as YapToText, which runs Whisper Large v3 Turbo with no setup.
How much memory does Whisper need?
It depends on the model. whisper.cpp lists about 273 MB for tiny, 852 MB for small, 2.1 GB for medium and 3.9 GB for large, and says quantized models need less. It gives no memory figure for turbo.
What is the best local Whisper model?
Large v3 Turbo is where I would start, and its 5-bit file is the one YapToText ships: 547 MiB in whisper.cpp's list. If memory is tight, small or base use far less, and OpenAI says the English-only versions tend to do better for English at the smallest sizes.
How big is the Whisper Large v3 Turbo model?
1.5 GiB in whisper.cpp's model list, or 547 MiB as the 5-bit q5_0 file, which is 574 MB counted in decimal units. OpenAI lists turbo at 809 million parameters.
Does local Whisper need the internet?
Only to download the model once. After that, whisper.cpp and YapToText transcribe with no connection, and the App Store build of YapToText ships with the model inside, so it skips even that download.
Can I use a whisper.cpp model in YapToText?
Yes. Under Your own models on the AI Models page, Add Model takes any whisper.cpp GGML .bin file, and the app tells what it is from the file itself.