Get started
Choosing a model
Find a speech engine that fits your hardware, language, and workflow.
On this page
A speech engine turns audio into text. AI cleanup uses a separate text model afterward. Choose the speech engine for your computer, language, and recording workflow first.
Start with the default#
For your first recording, keep the app's default: Parakeet on a new Windows x64 install or Local Whisper on other platforms. Your existing choice is preserved when you update.
After that, test alternatives on a short recording you know well. Compare the words, processing time, and memory use on your own computer; a larger model does not always give a better result for your recordings.
Compare speech engines#
| Engine | Where it runs | Supported workflows |
|---|---|---|
| Local Whisper | CPU on all platforms; NVIDIA CUDA on Windows/Linux | Dictation with preview, uploads, meetings |
| Parakeet | Windows/Linux CPU or NVIDIA GPU; Apple Silicon CPU or Metal; Intel Mac CPU from source | Dictation with preview, uploads, meeting chunks |
| Parakeet MLX | Apple Silicon, macOS 14+; Apple GPU or CPU | Dictation with preview, uploads, meeting chunks; experimental |
| Qwen3-ASR | Windows CPU or NVIDIA GPU; Apple Silicon CPU or MPS | Dictation and uploads |
| Nemotron Streaming | Windows/Linux CPU or NVIDIA GPU; Apple Silicon CPU or Metal; Intel Mac CPU from source | Dictation with preview, uploads, meetings with native preview |
| Moonshine | Windows CPU; Apple Silicon CPU on macOS 15+ | English dictation, uploads, meetings with native preview |
| OpenAI API | Cloud; API key and internet required | Dictation and uploads |
| Remote computer | The paired host's selected engine | Dictation, uploads, meetings with a supported host model |
Parakeet MLX remains experimental in this release. Orukeet is an optional community adaptation of Parakeet; Apple Silicon validation for Orukeet is pending. Evaluate these models on your recordings before relying on them.
Check language support#
Language choices depend on the engine. Optional local engines expose English, Russian, Spanish, French, Portuguese, Mandarin, and Auto where supported.
- Moonshine supports English only.
- Parakeet and Orukeet omit Mandarin from these presets.
- Qwen3-ASR supports all of these presets.
- Nemotron includes Mandarin in a broader coverage tier; accuracy may vary.
- Auto detects languages supported by the selected model.
Language selection transcribes speech; it does not translate it into another language. A remote computer offers the choices its host engine supports.
Download models and runtimes#
Use Settings → Downloads to see model details, download sizes, hardware estimates, required components, publishers, and licenses. Download both the model weights and any runtime the engine needs.
Local Whisper runs on CPU on macOS. On Windows or Linux x86_64, its optional GPU Acceleration component requires an NVIDIA driver providing CUDA 12 (525+); the CUDA Toolkit is not needed. Other engines have their own runtimes.
On Apple Silicon, Auto uses the Apple GPU where the engine supports it. Parakeet and Nemotron use NVIDIA Speech GPU (Metal); Qwen uses PyTorch MPS. Parakeet MLX has its own runtime and an approximately 2.5 GB model download. Moonshine requires macOS 15.
Use a cloud or remote engine#
For cloud transcription, choose the OpenAI API and add a key in Settings → API keys. Audio goes to the cloud provider, and provider charges may apply.
To use another computer's hardware, follow Remote engines. Audio goes to the host; it is not transcribed on the client computer.
See Privacy & offline for the network behavior of each option.
Add a custom Whisper model#
Choose Add custom models… from Downloads, Dictation → Voice model, or Meeting Mode → Voice & speakers. Add a local folder or discover a compatible Hugging Face repository and optional subfolder.
Custom models must use CTranslate2 Whisper format with model.bin, a valid config.json, and tokenizer.json. Original PyTorch weights need conversion first. Removing an entry leaves the source files in place.
Get help from the community, or tell us what went wrong.