Scrieb
Windows 10/11 · Guide · Updated Sep 2026

Whisper for Windows: Every Way to Run It

OpenAI's Whisper runs well on a Windows PC. The question is how much setup you want. This guide covers all five routes, from install-and-done apps to the exact command-line steps, with the GPU setup and the errors everyone hits.

Everything here runs locally. Your audio stays on the PC with all of these options.

Skip the setup: download Scrieb

What you need

Windows 10 or 11

64-bit. ARM64 machines (Snapdragon) run the apps; the Python route expects x64.

RAM

8 GB minimum. 16 GB is comfortable for the large model.

GPU, optional but decisive

NVIDIA RTX 20-series or newer: a one-hour file in 4–12 minutes. CPU only: about an hour for an hour.

Disk

1–3 GB per model, downloaded once. The audio files and transcripts stay wherever you put them.

The five ways, easiest first

SetupSpeaker labelsGPUPrice
1. Scrieb (app)Installer✓NVIDIA, AMD, IntelPaid, free trial
2. Buzz (app)Installer✗NVIDIAFree
3. WhisperDesktop (GUI)Unzip + model file✗Any (DirectCompute)Free
4. whisper.cpp (binaries)Unzip + terminal✗NVIDIA, CPU-optimisedFree
5. Official CLI (Python)Python + pip + ffmpeg✗ (WhisperX adds it)NVIDIA via CUDAFree

1. Scrieb: Whisper with no setup, plus speaker labels

Scrieb is a desktop app for Windows and Mac that bundles Whisper, a local diarization model for speaker labels, and a small local model for summaries. Install, drop in an audio or video file, get a transcript that reads as a dialogue, export DOCX, PDF or TXT. It needs the internet once for the model download, then works offline. Here is what that looks like:

It is the paid option on this page. If you only transcribe single-voice recordings and don't mind a terminal, the free routes below are genuinely fine. If you transcribe interviews, meetings or calls and need who-said-what without touching Python, this is the one. Download the free trial.

2. Buzz: free app, no speaker labels

Open-source desktop app for Windows, Mac and Linux. Download the installer from the project's GitHub releases, open it, drag in a file, pick a model size, get text. It can use an NVIDIA GPU. Export is basic and there are no speaker labels, which is fine for dictation, lectures and memos.

3. WhisperDesktop: free GUI that uses any GPU

A Windows-only GUI around whisper.cpp, notable because it uses DirectCompute, so AMD and Intel GPUs accelerate it, not just NVIDIA. Download the zip from the Const-me/Whisper GitHub releases, download a ggml model file, point the app at it. Updates are infrequent, and there are no speaker labels, but on a PC without an NVIDIA card it is the fastest free option.

4. whisper.cpp: prebuilt binaries, runs well on CPU

The C++ port of Whisper, with Windows binaries on its GitHub releases page. Unzip, download a model, run from a terminal. It wants 16 kHz mono WAV input, so convert first:

ffmpeg -i interview.mp3 -ar 16000 -ac 1 interview.wav
whisper-cli.exe -m ggml-medium.bin -f interview.wav -l en -otxt

The CUDA build on the releases page uses NVIDIA GPUs; the standard build is tuned for CPUs and is the best choice for older laptops.

5. The official command line: step by step

The reference implementation from OpenAI. About fifteen minutes the first time.

  1. Install Python 3.10 or 3.11 from python.org. Tick "Add Python to PATH" in the installer.
  2. Install ffmpeg, which Whisper uses to read audio. In a terminal:
    winget install Gyan.FFmpeg
    Close and reopen the terminal afterwards so PATH updates.
  3. Install Whisper:
    pip install -U openai-whisper
  4. If you have an NVIDIA GPU, install the CUDA build of PyTorch. pip installs the CPU build by default, which is the single most common reason Whisper is slow. Use the command from the selector at pytorch.org for your CUDA version; it looks like:
    pip install torch --index-url https://download.pytorch.org/whl/cu121
    Then check: python -c "import torch; print(torch.cuda.is_available())" must print True.
  5. Run it:
    whisper interview.mp3 --model medium --language en --output_format txt
    The first run downloads the model, 1.5 GB for medium, 3 GB for large. Output goes next to the input file as .txt, or .srt and .vtt with other output formats.
  6. Want speaker labels? Install WhisperX instead (pip install whisperx); it adds diarization but needs a Hugging Face token for the speaker model and the same CUDA setup.

A faster drop-in for the same workflow is faster-whisper, which uses less memory and runs the large model on smaller GPUs.

Errors everyone hits

"ffmpeg not found" / FileNotFoundError

ffmpeg is not on PATH. Install it with winget, then open a new terminal window; the old one keeps the old PATH.

It runs, but very slowly

torch.cuda.is_available() returns False: the CPU build of PyTorch is installed. Uninstall torch and install the CUDA build from pytorch.org.

"FP16 is not supported on CPU; using FP32 instead"

Harmless. It just means no GPU was found. Add --fp16 False to silence it, or fix the GPU setup.

First run hangs for minutes

It is downloading the model to %USERPROFILE%\.cache\whisper. Wait, or pre-download on a fast connection.

Invented sentences during silence or music

Whisper hallucinates on non-speech. Trim silence first, or use an app with voice-activity detection in front of the model.

Out of memory on large

Use medium, or large-v3-turbo, or faster-whisper with int8 compute. The large model needs about 10 GB of VRAM in the official build.

Which model on which hardware

ModelAccuracyCPU-only PCNVIDIA GPU
tiny / baseLowFast, drafts onlyInstant
smallGoodUsable, ~2× real timeVery fast
mediumVery goodAbout real time; the CPU defaultFast
large-v3BestToo slow for daily useThe default with a GPU
large-v3-turboNear large-v3SlowFastest high-accuracy option

Pricing

Whisper with speaker labels and no setup. One licence for Windows and Mac, no minute limits. Free trial first.

Monthly
$24.99
/month
Get started
Best Value
Yearly
$169.99
/year
Get started
6 Months
$99.99
/6 months
Get started

Frequently asked questions

Does Whisper work on Windows?
Yes. OpenAI's Whisper runs on Windows 10 and 11 through Python, through the whisper.cpp binaries, or inside desktop apps that bundle it. Everything on this page runs locally; the audio never leaves the PC.
Does Whisper need a GPU on Windows?
No, but it helps a lot. On CPU alone the medium model runs at roughly real time: a one-hour file takes about an hour. With an NVIDIA RTX 20-series or newer it runs 5 to 15 times faster. AMD and Intel GPUs work with whisper.cpp and with some apps; the official Python version is NVIDIA-only for GPU acceleration.
Is there a Whisper GUI for Windows?
Several. Scrieb is a full desktop app with speaker labels and DOCX export; Buzz is a free open-source app; WhisperDesktop is a free GUI around whisper.cpp that also uses AMD and Intel GPUs. None of the free GUIs label speakers.
Is Whisper free?
The model is open source and free. Running it costs nothing beyond your hardware. Paid apps like Scrieb charge for the interface, speaker identification, summaries and support, not for Whisper itself.
Which Whisper model should I use on Windows?
large-v3 if you have an NVIDIA GPU; medium on CPU-only machines; small for quick drafts of clean English audio. The turbo variant of large-v3 is a good middle ground on GPUs.
Can Whisper identify speakers?
Not by itself. Whisper produces text only. Speaker labels need a separate diarization model. On Windows, Scrieb runs that step locally; WhisperX does it from the command line; Buzz, WhisperDesktop and the official CLI don't label speakers.
Why is Whisper so slow on my PC?
Almost always one of three things: it is running on CPU because the CPU-only build of PyTorch was installed, the model is larger than the machine can handle, or the file is being processed in a single pass with a large model on little RAM. Check torch.cuda.is_available() first.

Related pages

→ Free Offline Transcription for Windows→ MacWhisper for Windows→ Offline Transcription for Windows→ What Is Whisper AI?→ What Is Speaker Diarization?