Parakeet vs Whisper for dictation on a Mac: NVIDIA's model against OpenAI's
By Ryleigh Newman · Published · Updated
Parakeet vs Whisper comes down to languages against speed: NVIDIA's Parakeet TDT 0.6B v3 is a 600-million-parameter speech model for 25 European languages that NVIDIA built for high throughput, and OpenAI's Whisper Large v3 Turbo is a multilingual model whose model card lists its many languages by name. Both can run locally on a Mac inside a dictation app, both add punctuation, and neither sends your voice anywhere when it runs on your machine.
I built YapToText, a free Mac dictation app that runs Whisper, and only Whisper. It does not offer Parakeet, so read what follows with that in mind. I built it because my hands make typing difficult and I dictate every day. Parakeet facts here come from NVIDIA's own model card, paper and blog, Whisper facts from OpenAI's own pages, and Mac app facts from each app's own site. No speed or accuracy number on this page is mine.
Download for free. No account, nothing locked, macOS 14 or later on Apple silicon. Source on GitHub.
On this page: 11 sections
Steps
- Write down five sentences you really say, with a name that usually comes out wrong and one sentence in each language you dictate.
- Dictate them in an app that offers Parakeet V3, such as Handy or VoiceInk, with any AI cleanup switched off, and count the words you would have to fix.
- Quit that app, then dictate the same sentences in YapToText with Raw Transcription, by pressing 2 while you talk, and count the fixes the same way.
- In YapToText's History, click Show pipeline details to see how long transcription took on your Mac, and compare it with the other app.
Two models, two makers
The two names get used like brand names, so it helps to pin down exactly which files people mean. Parakeet V3 and Whisper Large v3 Turbo are the two local models most Mac dictation apps now put in front of you.
- Parakeet TDT 0.6B v3. NVIDIA's model card on Hugging Face calls it a 600-million-parameter multilingual speech recognition model designed for high-throughput speech-to-text transcription. It extends Parakeet TDT 0.6B v2 from English to 25 European languages. TDT stands for Token-and-Duration Transducer, the decoder that sits behind a FastConformer encoder. NVIDIA releases it under the CC BY 4.0 license.
- Parakeet V2. The English-only model v3 grew out of. VoiceInk's docs suggest Parakeet V2 or Parakeet Unified when you only dictate English.
- Whisper Large v3 Turbo. OpenAI calls Whisper a general-purpose speech recognition model trained on a large dataset of diverse audio, and turbo an optimized version of large-v3 with faster transcription and a minimal degradation in accuracy. OpenAI lists turbo at 809 million parameters and released Whisper's code and weights under the MIT License.
- What YapToText runs. The stock Whisper Large v3 Turbo (Q5) file, 574 MB, through a copy of whisper.cpp built with Metal and Accelerate. The Whisper for Mac page covers that file in detail.
One thing the names hide: Parakeet's model card describes running it with NVIDIA's NeMo toolkit and lists NVIDIA GPU architectures, such as Ampere and Hopper, as its hardware. An Apple silicon Mac has no NVIDIA GPU, so a Mac app that offers Parakeet has to run it some other way, and its own docs say how. Whisper on a Mac usually means whisper.cpp, whose README calls Apple silicon a first-class citizen.
Side by side
Checked October 5, 2026, on NVIDIA's Parakeet TDT 0.6B v3 model card and OpenAI's Whisper Large v3 Turbo model card on Hugging Face, and October 3, 2026, on OpenAI's Whisper README and whisper.cpp's README.
| Parakeet TDT 0.6B v3 | Whisper Large v3 Turbo | |
|---|---|---|
| Maker | NVIDIA | OpenAI |
| Size | 600 million parameters | 809 million parameters |
| Languages | 25 European languages | Dozens, listed by name on the model card; performance varies widely by language, per OpenAI |
| Picking the language | Detects it automatically, with no prompting | Set by the app, or detected; YapToText offers 29 languages plus Auto-detect |
| Punctuation and capitals | Automatic, per the model card | Punctuates ordinary speech itself in YapToText |
| Timestamps | Word-level and segment-level, per the model card | Depends on the app; YapToText's file transcripts are plain text with none |
| Long audio | Up to 24 minutes in one pass with full attention on an A100 80GB GPU, or up to 3 hours with local attention | Depends on the app |
| License | CC BY 4.0 | MIT |
| Made to run on | NVIDIA's NeMo toolkit, on NVIDIA GPUs | Anything that runs it; whisper.cpp treats Apple silicon as first-class |
Two rows matter most for dictation. Languages decide whether Parakeet can hear you at all, and the runtime decides how fast either one feels on your Mac. Parakeet is the narrower, faster-built model; Whisper is the broader one.
Speed and accuracy: what the makers measured
The numbers people quote for Parakeet vs Whisper come from the makers and from leaderboards, so here is exactly what they measured. Every figure in this section is NVIDIA's or OpenAI's measurement, not mine.
Checked October 5, 2026, in NVIDIA's paper on arXiv and NVIDIA's blog.
- Accuracy, by NVIDIA. NVIDIA's paper, "Canary-1B-v2 & Parakeet-TDT-0.6B-v3", compares English speech recognition on the test sets of the Hugging Face Open ASR Leaderboard, such as LibriSpeech, AMI and Earnings-22. It reports an average word error rate of 6.32% for Parakeet TDT 0.6B v3 and 7.44% for Whisper large-v3. Lower is better: it is the share of words the model got wrong.
- Speed, by NVIDIA. The same table gives Parakeet an RTFx of 3,332.74, far higher than Whisper large-v3's. NVIDIA's blog defines throughput on that leaderboard as the duration of audio transcribed divided by the computation time, and says Parakeet v3 has the highest throughput of the multilingual models there.
- Turbo, by OpenAI. The paper's row is Whisper large-v3, not turbo. OpenAI rates turbo at about 8 times the speed of large with a minimal degradation in accuracy, measured on an A100 GPU, not on a Mac.
Read those numbers with four limits in mind. They are English test sets, so they say nothing about your Polish or your Japanese. They were measured on NVIDIA's hardware, while a Mac app runs either model on Apple silicon, often from a converted file, so the speed ratio will not carry over as it is. They compare large-v3, not the 5-bit turbo file YapToText ships. And a dictation is a few seconds of one voice, not a benchmark recording.
On made-up words, NVIDIA's paper says non-speech audio was added to the training data to reduce hallucinations. OpenAI's Whisper model card says Whisper may output text that was not actually spoken, and the Whisper hallucinations guide covers why and what YapToText does about it. I found no source that measures the two side by side on that, so I make no claim either way.

Which languages each one hears
For most people the choice is settled here, before speed comes into it. If Parakeet does not list your language, it is not your model.
NVIDIA's model card lists these 25: Bulgarian, Croatian, Czech, Danish, Dutch, English, Estonian, Finnish, French, German, Greek, Hungarian, Italian, Latvian, Lithuanian, Maltese, Polish, Portuguese, Romanian, Russian, Slovak, Slovenian, Spanish, Swedish and Ukrainian. Here is how that lines up with the 29 languages in YapToText's Language picker:
- In both. Czech, Danish, Dutch, English, Finnish, French, German, Greek, Hungarian, Italian, Polish, Portuguese, Romanian, Russian, Slovak, Spanish, Swedish and Ukrainian.
- Only in YapToText's picker. Arabic, Chinese, Hebrew, Hindi, Indonesian, Japanese, Korean, Norwegian Bokmål, Thai, Turkish and Vietnamese.
- Only in Parakeet v3. Bulgarian, Croatian, Estonian, Latvian, Lithuanian, Maltese and Slovenian. Whisper's model card lists all seven, but YapToText's picker does not, and I only claim what the picker lists.
Two more differences show up when you dictate in more than one language. Parakeet detects the language on its own. YapToText sets one language per dictation, from the picker or Auto-detect, and a change applies from the next dictation. The languages guide shows how to tie a language to an app or a shortcut.
Mac apps that run Parakeet, and where YapToText stands
Many Mac dictation apps now offer Parakeet next to Whisper, and the site's own comparison pages already track which ones. If you want Parakeet on a Mac today, you get it through one of these apps, not from YapToText.
Checked October 3, 2026, on each app's own site, docs or repository.
| App | Parakeet | Whisper | Price |
|---|---|---|---|
| VoiceInk | Parakeet V3 by default at setup; Parakeet V2 or Parakeet Unified for English only | Yes | 7-day trial, then $25 once for one Mac |
| Spokenly | Local Parakeet models, free | Yes, free | Free with local models; Pro $9.99 a month |
| superwhisper | On Pro only | Yes, on the free plan | Free plan; Pro $8.49 a month |
| Handy | Parakeet V3 | Small, Medium, Turbo and Large | Free, MIT licensed |
| OpenWhispr | Yes, among its local models | Yes | Free local dictation |
| FluidVoice | Yes, plus Nemotron | Yes | Free, GPL-3.0 |
| YapToText | No | Large v3 Turbo (Q5) inside, ten more in the library | Free |
The detail pages are the VoiceInk comparison, the Spokenly comparison, the superwhisper comparison and the roundup of open-source dictation apps, which covers Handy, OpenWhispr and FluidVoice.
On the YapToText side, the model choice lives on the AI Models page. The Dictation Model Library has ten Whisper models besides the bundled one, and Add Model, under Your own models, takes a whisper.cpp .bin or a llama.cpp .gguf file and rejects anything else. So a Parakeet file will not load into it. The AI models guide lists every model with the size the app shows.

Who should pick which
Here is how I would choose, honestly, even where the answer is not my app. Pick the model by your languages first and your Mac's speed second.
An app with Parakeet may suit you if
- You only dictate in European languages. Above all if one of them is Bulgarian, Croatian, Estonian, Latvian, Lithuanian, Maltese or Slovenian, which YapToText's picker does not list.
- Speed is the thing you notice most. NVIDIA built it for throughput, and its paper's numbers point that way. Measure it on your own Mac before you believe it.
- You switch between European languages a lot. Parakeet detects the language itself, with no picker to change.
- You want word timestamps on files. The model card lists word-level timestamps, though whether you see them depends on the app.
Whisper may suit you if
- You dictate outside Europe's languages. Japanese, Chinese, Korean, Hindi, Arabic, Hebrew, Thai, Vietnamese, Indonesian, Turkish or Norwegian Bokmål are in YapToText's picker and not in Parakeet v3's list.
- You want to choose a size. Whisper comes from Tiny to Large, so you can trade accuracy for battery and memory. On the Energy page, Switch models with the power source can even run a lighter model on battery.
- You want it with nothing to set up. YapToText's App Store build ships Whisper inside and works offline from the first launch.
- You want to run it yourself. whisper.cpp lists Intel and Apple silicon Macs as supported, and the local Whisper guide has the commands.
Translation does not decide this for dictation. YapToText only asks Whisper to transcribe, and OpenAI says turbo is not trained for translation anyway. If you are weighing Whisper against the Mac's own engine instead, Apple Dictation vs Whisper covers that pairing.
Test both on your own voice
The only fair test is your own voice, your own sentences and your own Mac. It takes about fifteen minutes. Compare raw against raw, so no cleanup model gets the credit or the blame.
Step 1: Write down five sentences you really say
- You: Pick sentences you send every week, with at least one name that usually comes out wrong, and one in each language you dictate.
- Check: Read them aloud once so the pace is natural.
Step 2: Dictate them with Parakeet
- You: In an app that offers Parakeet V3, such as Handy, which is free, or VoiceInk, whose setup downloads Parakeet V3 by default, choose Parakeet V3 and switch off any AI cleanup the app offers. Say the five sentences into a note.
- Check: Count the words you would have to fix, then quit that app so only one app answers your dictation key.
Step 3: Dictate them with Whisper in YapToText
- You: Get YapToText from the Mac App Store. Click into the note, tap Right Command, press 2 for Raw Transcription while you talk, say the same sentences and tap again.
- YapToText: Types exactly what Whisper heard, with no cleanup model involved.
- Check: Count the fixes the same way. The first dictation after launch waits a moment while the model loads, so time the second one.
Step 4: Compare the speed on your Mac
- You: In YapToText, open History, find the dictation and click Show pipeline details.
- YapToText: Splits the stop-to-text time into transcription, AI cleanup and delivery.
- Check: The transcription line is Whisper's time on your Mac. Hold it up against how long the other app took to type the same sentences.
If Whisper keeps missing a name, add it on the Dictionaries page before you decide: each entry primes Whisper with the word as well as fixing its spelling.
What each one cannot do
Parakeet TDT 0.6B v3
- Hear languages outside its 25. Japanese, Chinese, Arabic, Hindi and the rest are not on its list.
- Run on a Mac by itself. Its model card describes NVIDIA's NeMo toolkit on NVIDIA GPUs, so on a Mac it comes through an app.
- Run inside YapToText. YapToText runs Whisper models only, and Add Model rejects anything that is not a whisper.cpp or llama.cpp file.
Whisper Large v3 Turbo
- Hear every language equally. OpenAI says performance varies widely depending on the language.
- Translate. OpenAI says turbo is not trained for translation.
- Promise it never invents a word. OpenAI's model card says it may output text that was not spoken, which is why YapToText guards against it.
- Run YapToText on an Intel Mac. YapToText needs Apple silicon and macOS 14 or later, though whisper.cpp itself supports Intel Macs.
Neither model is the whole answer; the app around it decides how it feels to dictate. I have not run either model through a benchmark, and I do not publish accuracy percentages for any model.
The short version
- Parakeet. NVIDIA's 600-million-parameter model for 25 European languages, built for throughput, CC BY 4.0.
- Whisper. OpenAI's MIT-licensed model, dozens of languages on the turbo model card, sizes from Tiny to Large.
- The numbers. NVIDIA's paper puts Parakeet ahead of Whisper large-v3 on English test sets and far ahead on throughput; those are NVIDIA's measurements, not mine.
- On a Mac. VoiceInk, Spokenly, superwhisper Pro, Handy, OpenWhispr and FluidVoice offer Parakeet; YapToText runs Whisper only.
European languages and raw speed point to Parakeet; every other language, and a free app with the model inside, point to Whisper.
If you tested both on your own voice and something here does not match what you found, open an issue on GitHub. Every one gets read.
Download for free. No account, nothing locked, macOS 14 or later on Apple silicon. Source on GitHub.
Related
Sources
- https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3 (checked 2026-10-05)
- https://arxiv.org/abs/2509.14128 (checked 2026-10-05)
- https://blogs.nvidia.com/blog/speech-ai-dataset-models/ (checked 2026-10-05)
- https://huggingface.co/openai/whisper-large-v3-turbo (checked 2026-10-05)
- https://github.com/openai/whisper (checked 2026-10-03)
- https://github.com/openai/whisper/blob/main/model-card.md (checked 2026-10-05)
- https://github.com/ggml-org/whisper.cpp (checked 2026-10-03)
- https://tryvoiceink.com/pricing (checked 2026-10-03)
- https://tryvoiceink.com/docs/installation (checked 2026-10-03)
- https://tryvoiceink.com/docs/recommended-models (checked 2026-10-03)
- https://spokenly.app (checked 2026-10-03)
- https://spokenly.app/pricing (checked 2026-10-03)
- https://superwhisper.com (checked 2026-10-03)
- https://superwhisper.com/docs/billing/plans (checked 2026-10-03)
- https://github.com/cjpais/Handy (checked 2026-10-03)
- https://github.com/OpenWhispr/openwhispr (checked 2026-10-03)
- https://openwhispr.com/pricing (checked 2026-10-03)
- https://github.com/altic-dev/FluidVoice (checked 2026-10-03)
- https://github.com/ryleighnewman/YapToText (checked 2026-10-03)
- https://apps.apple.com/us/app/yaptotext/id6786382289 (checked 2026-10-03)
Questions
Is Parakeet better than Whisper?
It depends on your language. NVIDIA's own paper reports a lower average word error rate and far higher throughput for Parakeet TDT 0.6B v3 than for Whisper large-v3 on English test sets, but Parakeet v3 covers 25 European languages, while Whisper's turbo model card lists dozens more. Test both on your own sentences.
Can I use Parakeet on a Mac?
Yes, through a dictation app. VoiceInk, Spokenly, Handy, OpenWhispr and FluidVoice offer Parakeet models, and superwhisper offers them on Pro. YapToText does not; it runs Whisper only.
What languages does Parakeet V3 support?
NVIDIA's model card lists 25 European languages: Bulgarian, Croatian, Czech, Danish, Dutch, English, Estonian, Finnish, French, German, Greek, Hungarian, Italian, Latvian, Lithuanian, Maltese, Polish, Portuguese, Romanian, Russian, Slovak, Slovenian, Spanish, Swedish and Ukrainian. It detects the language automatically.
Is Parakeet faster than Whisper?
By NVIDIA's measurements, yes: its paper gives Parakeet TDT 0.6B v3 an RTFx of 3,332.74, far higher than Whisper large-v3's, measured by NVIDIA on its own hardware. How either feels on your Mac depends on the app, so time a few dictations yourself.
Does YapToText support Parakeet?
No. YapToText runs Whisper Large v3 Turbo (Q5) by default, offers ten more Whisper models in its library, and its Add Model button takes only whisper.cpp and llama.cpp files.
Is Parakeet free and open?
NVIDIA publishes Parakeet TDT 0.6B v3 on Hugging Face under the CC BY 4.0 license. Whisper's code and weights are under the MIT License.