Guides

Whisper Large v3 Turbo on a Mac: what it is, how it compares, and how to run it free

Whisper Large v3 Turbo is OpenAI's faster version of its large-v3 speech model, and you can run it free on a Mac, on device, either inside an app or with whisper.cpp in Terminal. OpenAI's README calls turbo an optimized version of large-v3 that offers faster transcription speed with a minimal degradation in accuracy, lists it at 809 million parameters against 1,550 million for large, and says it is not trained for translation.

It is also the model I chose for YapToText, the free dictation app I built because my hands make typing difficult: the App Store download ships Whisper Large v3 Turbo (Q5), the stock 5-bit whisper.cpp file, so weigh what I say about the app with that in mind. This page sticks to turbo itself. For every Whisper size and the full whisper.cpp setup, the local Whisper guide goes further, and the AI models guide covers YapToText's model library.

Download YapToText on the Mac App StoreDownload for free. No account, nothing locked, macOS 14 or later on Apple silicon. Source on GitHub.

On this page: 11 sections
  1. Steps
  2. What turbo is, in OpenAI's words
  3. Turbo vs large-v3 and the smaller models
  4. Q5, Q8 or the full file
  5. Run turbo in YapToText, with no Terminal
    1. Step 1: Install it
    2. Step 2: Grant two permissions
    3. Step 3: Dictate
  6. Run turbo yourself with whisper.cpp
  7. See how turbo runs on your own Mac
  8. If turbo mishears you or feels slow
  9. What turbo cannot do
  10. The short version
  11. Questions

Steps

  1. Install YapToText from the Mac App Store; Whisper Large v3 Turbo (Q5) is inside.
  2. Allow the microphone, then turn on Accessibility for YapToText in System Settings, Privacy & Security.
  3. Check that the AI Models page lists Whisper Large v3 Turbo (Q5) under Your defaults.
  4. Click where you want the words, tap Right Command, talk, and tap it again.
  5. Open the dictation in History and click Show pipeline details to see how long transcription took on your Mac.

What turbo is, in OpenAI's words

Turbo is not a separate family of models. It is large-v3, trimmed so it runs faster. Every fact in this section comes from OpenAI's Whisper README on GitHub and OpenAI's model card on Hugging Face. Turbo is large-v3 with most of its decoder taken out, traded for speed.

  • What changed. OpenAI's model card on Hugging Face says the decoding layers were cut from 32 to 4. The README sums it up as an optimized version of large-v3 with faster transcription and a minimal degradation in accuracy.
  • Size. 809 million parameters, against 1,550 million for large. That is roughly half the weights.
  • Speed. The README rates turbo at about 8 times the speed of large. OpenAI measured that by transcribing English speech on an A100, a data-center GPU, and says real-world speed may vary significantly depending on the language, the speaking speed and the hardware. It is not a Mac figure.
  • Memory in OpenAI's package. About 6 GB of required VRAM for turbo and about 10 GB for large. Those numbers are for OpenAI's own Python package, not for whisper.cpp or any Mac app.
  • Translation. OpenAI says turbo is not trained for translation tasks and returns the original language even if you ask it to translate. For speech to English translation, it points you to the multilingual tiny, base, small, medium or large models instead.
  • The default. In OpenAI's own package, the default setting selects turbo, and the README says it works well for transcribing English.
  • The license. OpenAI released Whisper's code and model weights under the MIT License, which is why free apps can ship it.

Two cautions carry over from large-v3 unchanged. OpenAI says Whisper's performance varies widely depending on the language, and its model card says Whisper may output text that was not actually spoken. The Whisper hallucinations guide covers that second one, and what YapToText does about it.

Turbo vs large-v3 and the smaller models

Turbo vs large-v3 comes down to one question: is the slower model worth it? OpenAI's own wording is the honest answer: turbo is faster, with a minimal degradation in accuracy. I do not publish accuracy numbers for any model, and I have not benchmarked these files against each other. For dictation, turbo's speed is the point; large-v3 is for when you want the full decoder or need translation.

Checked October 3, 2026, in OpenAI's Whisper README and in whisper.cpp's README and model list. The last column is the size YapToText's Dictation Model Library shows.

ModelParameters (OpenAI)File on disk (whisper.cpp)Size in YapToText
small244 M466 MiB466 MB
medium769 M1.5 GiB1.5 GB
large-v31550 M2.9 GiB3.0 GB
large-v3, 5-bit (q5_0)Quantized large-v31.1 GiB1.1 GB
large-v3-turbo809 M1.5 GiB1.5 GB
large-v3-turbo, 8-bit (Q8)Quantized turboNot in whisper.cpp's table834 MB
large-v3-turbo, 5-bit (q5_0)Quantized turbo547 MiB574 MB, bundled

MiB counts in steps of 1,024, so the 547 MiB turbo file in whisper.cpp's list is the same file YapToText shows as 574 MB. Here is how I would read the table:

  • Turbo or large-v3. Turbo has about half the parameters and, by OpenAI's measurement on its own hardware, about 8 times the speed of large. YapToText's library calls Whisper Large v3 the largest and slowest model it offers.
  • Turbo or medium. The full turbo file and medium are the same size on disk in whisper.cpp's list, and turbo is built from large-v3.
  • Turbo or something small. If memory is tight and you only speak English, OpenAI says the English-only .en models tend to do better, especially tiny.en and base.en.
  • Translation. If you need speech translated to English, turbo is the wrong pick, by OpenAI's own note.

The Parakeet vs Whisper comparison sets turbo against NVIDIA's model, and Apple Dictation vs Whisper sets it against the engine built into macOS.

Q5, Q8 or the full file

On a Mac you rarely run the original turbo weights. whisper.cpp converts Whisper into its own ggml format, and it also publishes quantized files, which end in names like q5_0. whisper.cpp says quantized models need less memory and disk space and, depending on the hardware, can be processed more efficiently. The 5-bit turbo file is about a third of the full turbo file on disk.

  • Whisper Large v3 Turbo (Q5). 574 MB in YapToText, the file ggml-large-v3-turbo-q5_0.bin, and the default since 1.3.1, the release whose notes say dictation became over twice as fast and the app about a gigabyte smaller.
  • Whisper Large v3 Turbo (Q8). 834 MB, one click away in YapToText's Dictation Model Library.
  • Whisper Large v3 Turbo. The full file, 1.5 GB.

I picked the 5-bit file because, on my own dictations, it heard almost exactly what the full model heard, with a third of the memory. That is one person's dictations, not a benchmark. Its SHA-256 matches the published file on Hugging Face, so it is the stock whisper.cpp build, which I have not fine-tuned, retrained or modified.

whisper.cpp's memory table covers only the five original sizes, from about 273 MB for tiny to about 3.9 GB for large, and gives no figure for turbo or for the quantized files. So nobody can quote you a whisper.cpp memory number for turbo from that table. Your own Mac is the test, and the next section shows where to look.

The AI Models page in YapToText: Whisper Large v3 Turbo (Q5) and Phi-3.5 Mini Instruct marked Included, model library below.
Figure 1. Whisper Large v3 Turbo (Q5) sits under Your defaults at 574 MB, and the Q8 and full turbo files wait in the Dictation Model Library below it.

Run turbo in YapToText, with no Terminal

YapToText runs the 5-bit turbo file through a copy of whisper.cpp built into the app with Metal and Accelerate, and types what you say at the cursor in any app. It is free on the Mac App Store, with no account, for an Apple silicon Mac with macOS 14 or later. The model is inside the download, so there is nothing to fetch and your voice never leaves your Mac.

Step 1: Install it

  • You: Get YapToText from the Mac App Store. The download is about 2.9 GB, because both models are inside.
  • YapToText: Arrives ready to dictate offline, with Whisper Large v3 Turbo (Q5) as the dictation model.
  • Check: The AI Models page lists Whisper Large v3 Turbo (Q5) under Your defaults, marked Included.

Step 2: Grant two permissions

  • You: Allow the microphone when macOS asks, then press Grant next to Accessibility on the Home page and switch YapToText on in System Settings, Privacy & Security, Accessibility.
  • macOS: Asks for the microphone once. Accessibility is a switch only you can turn on.
  • Check: The Get set up card leaves the Home page once both are on.

Step 3: Dictate

  • You: Click where the words go, tap Right Command, talk, and tap it again.
  • YapToText: Runs turbo on your Mac and pastes the result at your cursor.
  • Check: The words are there. The first dictation after launch can take a moment while the model loads.

To try the Q8 or full turbo file, or Whisper Large v3, press the download button on its row in the Dictation Model Library, then click the circle at the start of the row once it says Installed. If the model you picked is not on disk yet, another downloaded Whisper model stands in, so dictation keeps working while it downloads. The Homebrew build ships without models, so on that build, download Whisper Large v3 Turbo (Q5) before your first dictation; the install guide covers both routes.

Run turbo yourself with whisper.cpp

If you want the model without an app, whisper.cpp is the engine to use on a Mac. Its README calls Apple silicon a first-class citizen, optimized through ARM NEON, Accelerate, Metal and Core ML, and lists Intel and Arm Macs as supported. These are the README's own commands, with the model set to the 5-bit turbo file. whisper.cpp gives you a transcript of an audio file in Terminal, not words typed into your apps.

Terminal: Get the code, download the 5-bit turbo model, build, and transcribe the bundled sample.

git clone https://github.com/ggml-org/whisper.cpp.git
cd whisper.cpp
sh ./models/download-ggml-model.sh large-v3-turbo-q5_0
cmake -B build
cmake --build build -j --config Release
./build/bin/whisper-cli -m models/ggml-large-v3-turbo-q5_0.bin -f samples/jfk.wav

What happens: The script downloads ggml-large-v3-turbo-q5_0.bin from Hugging Face, CMake builds whisper-cli, and whisper-cli prints the sample's transcript in Terminal. Swap large-v3-turbo-q5_0 for large-v3-turbo to get the full file.

Two things the README adds: whisper-cli currently runs only with 16-bit WAV files, so convert other audio with ffmpeg first, and whisper-stream can transcribe your microphone live if you install SDL2, though the README calls it a naive example. The local Whisper guide has the ffmpeg command and each step explained.

Terminal: Or use OpenAI's own Python package, where turbo is the default model.

pip install -U openai-whisper
whisper audio.mp3 --model turbo

What happens: pip installs the whisper command, which transcribes audio.mp3 with turbo. OpenAI's README says the code is expected to work with Python 3.8 to 3.11 and needs ffmpeg.

See how turbo runs on your own Mac

OpenAI's speed ratio was measured on an A100, and whisper.cpp publishes no memory figure for turbo, so the only numbers that describe your Mac are the ones your Mac produces. Every dictation in YapToText's History carries its own timing, so you can see turbo's speed instead of trusting a chart.

  • Timing. Open any dictation in History and click Show pipeline details: the stop-to-text time is split into transcription, AI cleanup and delivery. Transcription is turbo's share.
  • Your Mac. The This Mac card on the Energy page shows your chip and memory.
  • Memory it gives back. Keep models in memory, on the Energy page, decides how long a model stays loaded after a dictation: 2 minutes after use by default. The next dictation after an unload takes a few extra seconds.
  • Room for macOS. When macOS runs short of memory, YapToText drops its loaded models at once and waits five minutes before warming them again.
  • A lighter model on battery. Switch models with the power source, off by default, gives the dictation model a plugged-in and an on-battery choice, so turbo can run on power and a smaller model on battery.
A YapToText History entry with pipeline details open: raw transcript, cleaned transcript, mode, delivery, target app and timings.
Figure 2. Show pipeline details splits each dictation's stop-to-text time into transcription, AI cleanup and delivery, which is turbo's speed on your Mac, not OpenAI's A100.

The energy guide explains each of these settings.

If turbo mishears you or feels slow

Most fixes cost less than a bigger download. Fix the microphone and the words before you change the model. Climb this ladder, smallest step first:

  1. Wait out the first dictation. The model loads once per launch, and again after it has been unloaded to save memory. The dictations after that are quick.
  2. Set your language. The Language picker on the AI Models page starts on your Mac's system language. The languages guide lists all 29, plus Auto-detect.
  3. Fix the microphone. Bluetooth microphones are phone quality, so pick the Mac's own microphone in the menu bar panel's Input picker.
  4. Teach it your words. Names and terms on the Dictionaries page prime Whisper, not just fix what it types. The dictionaries guide shows how.
  5. Try another turbo file or large-v3. The Q8 build of turbo, or Whisper Large v3, means a bigger download and more memory, and the star ratings on the AI Models page show the trade. If turbo is too slow on your Mac, go smaller instead.
  6. Tell me. Settings, About, Copy Diagnostics gathers your Mac model, versions, model setup and latency stats, never your words, for an issue on GitHub.

What turbo cannot do

  • Translate. OpenAI says turbo is not trained for translation and returns the original language even when asked to translate. YapToText only ever asks Whisper to transcribe; to translate afterward, select the text and ask Quick Edit for another language.
  • Hear every language equally. OpenAI says Whisper's performance varies widely depending on the language.
  • Promise it heard you. OpenAI's model card says Whisper may output text that was not actually spoken. Read anything important before you rely on it.
  • Carry its speed claim to your Mac. The 8 times figure is OpenAI's, from an A100. On a Mac, the file, the engine and the chip decide.
  • Run in YapToText on an Intel Mac. YapToText needs Apple silicon and macOS 14 or later, and nothing is claimed for Intel Macs. whisper.cpp lists Intel Macs as supported.

To check that a setup is really local, the offline dictation page has a one-minute test: turn Wi-Fi off and dictate. Local changes where your voice goes, not how well it is heard.

The short version

  • What it is. OpenAI's optimized large-v3: 809 million parameters, decoder cut from 32 layers to 4, faster with a minimal degradation in accuracy.
  • Turbo vs large-v3. Turbo for speed; large-v3 when you want the full model or need translation.
  • In an app. YapToText ships the 5-bit turbo file, 574 MB, and types into any app, free and on device.
  • In Terminal. whisper.cpp's download script fetches large-v3-turbo-q5_0, and whisper-cli transcribes WAV files.

Pick turbo for dictation, then let your own History tell you how fast it is on your Mac.

If something here does not match what you see, open an issue on GitHub. Every one gets read.

Download YapToText on the Mac App StoreDownload for free. No account, nothing locked, macOS 14 or later on Apple silicon. Source on GitHub.

About the author. Ryleigh Newman is a developer whose hands make typing difficult and who dictates every day with YapToText, the app this site describes. Facts about other apps come from their own sites, with the date they were read. More about the project.

Sources

Questions

Is Whisper Large v3 Turbo better than Large v3?

It is faster, not more accurate. OpenAI calls turbo an optimized version of large-v3 with faster transcription and a minimal degradation in accuracy, and rates it at about 8 times the speed of large, measured on an A100 GPU.

How big is Whisper Large v3 Turbo?

OpenAI lists it at 809 million parameters. whisper.cpp's file is 1.5 GiB, and its 5-bit q5_0 file is 547 MiB, which YapToText shows as 574 MB.

Can Whisper Large v3 Turbo translate?

No. OpenAI says turbo is not trained for translation and returns the original language even when asked to translate. For speech to English translation, OpenAI points to the multilingual tiny, base, small, medium or large models.

How do I run Whisper Large v3 Turbo on a Mac?

Install YapToText from the Mac App Store, which ships the 5-bit turbo file and types into any app. Or clone whisper.cpp, download large-v3-turbo-q5_0 with its script, build with CMake and run whisper-cli on a 16-bit WAV file.

How much memory does Whisper turbo need on a Mac?

OpenAI lists about 6 GB of VRAM for turbo in its Python package, and whisper.cpp gives no memory figure for turbo. YapToText uses the 5-bit file, which on my own dictations heard almost what the full model heard with a third of the memory.

Does Whisper turbo run on Apple silicon?

Yes. whisper.cpp's README calls Apple silicon a first-class citizen and uses Metal for the GPU. YapToText runs turbo through whisper.cpp on Apple silicon Macs with macOS 14 or later.