Skip to content

Get started

Choosing a model

Find a speech engine that fits your hardware, language, and workflow.

On this page

A speech engine turns audio into text. AI cleanup uses a separate text model afterward. Choose the speech engine for your computer, language, and recording workflow first.

Start with the default#

For your first recording, keep the app's default: Parakeet on a new Windows x64 install or Local Whisper on other platforms. Your existing choice is preserved when you update.

After that, test alternatives on a short recording you know well. Compare the words, processing time, and memory use on your own computer; a larger model does not always give a better result for your recordings.

Compare speech engines#

EngineWhere it runsSupported workflows
Local WhisperCPU on all platforms; NVIDIA CUDA on Windows/LinuxDictation with preview, uploads, meetings
ParakeetWindows/Linux CPU or NVIDIA GPU; Apple Silicon CPU or Metal; Intel Mac CPU from sourceDictation with preview, uploads, meeting chunks
Parakeet MLXApple Silicon, macOS 14+; Apple GPU or CPUDictation with preview, uploads, meeting chunks; experimental
Qwen3-ASRWindows CPU or NVIDIA GPU; Apple Silicon CPU or MPSDictation and uploads
Nemotron StreamingWindows/Linux CPU or NVIDIA GPU; Apple Silicon CPU or Metal; Intel Mac CPU from sourceDictation with preview, uploads, meetings with native preview
MoonshineWindows CPU; Apple Silicon CPU on macOS 15+English dictation, uploads, meetings with native preview
OpenAI APICloud; API key and internet requiredDictation and uploads
Remote computerThe paired host's selected engineDictation, uploads, meetings with a supported host model

Parakeet MLX remains experimental in this release. Orukeet is an optional community adaptation of Parakeet; Apple Silicon validation for Orukeet is pending. Evaluate these models on your recordings before relying on them.

Check language support#

Language choices depend on the engine. Optional local engines expose English, Russian, Spanish, French, Portuguese, Mandarin, and Auto where supported.

  • Moonshine supports English only.
  • Parakeet and Orukeet omit Mandarin from these presets.
  • Qwen3-ASR supports all of these presets.
  • Nemotron includes Mandarin in a broader coverage tier; accuracy may vary.
  • Auto detects languages supported by the selected model.

Language selection transcribes speech; it does not translate it into another language. A remote computer offers the choices its host engine supports.

Download models and runtimes#

Use Settings → Downloads to see model details, download sizes, hardware estimates, required components, publishers, and licenses. Download both the model weights and any runtime the engine needs.

Local Whisper runs on CPU on macOS. On Windows or Linux x86_64, its optional GPU Acceleration component requires an NVIDIA driver providing CUDA 12 (525+); the CUDA Toolkit is not needed. Other engines have their own runtimes.

On Apple Silicon, Auto uses the Apple GPU where the engine supports it. Parakeet and Nemotron use NVIDIA Speech GPU (Metal); Qwen uses PyTorch MPS. Parakeet MLX has its own runtime and an approximately 2.5 GB model download. Moonshine requires macOS 15.

Use a cloud or remote engine#

For cloud transcription, choose the OpenAI API and add a key in Settings → API keys. Audio goes to the cloud provider, and provider charges may apply.

To use another computer's hardware, follow Remote engines. Audio goes to the host; it is not transcribed on the client computer.

See Privacy & offline for the network behavior of each option.

Add a custom Whisper model#

Choose Add custom models… from Downloads, Dictation → Voice model, or Meeting Mode → Voice & speakers. Add a local folder or discover a compatible Hugging Face repository and optional subfolder.

Custom models must use CTranslate2 Whisper format with model.bin, a valid config.json, and tokenizer.json. Original PyTorch weights need conversion first. Removing an entry leaves the source files in place.

Checked against OpenWhisper 2.6.14.App documentation
Still stuck?

Get help from the community, or tell us what went wrong.