Scrieb
Speaker Identification · Mac & Windows · Offline

Transcription with Speaker Identification – Without Uploading Anything

A transcript without speaker labels is a wall of text. Scrieb tells you who said what — automatically, in every plan, and entirely on your own computer.

Interviews, meetings, focus groups, calls: drop in the recording, get a labelled dialogue back, export it as DOCX or PDF. No cloud, no minute limits.

Download Scrieb

What speaker identification does

Speaker identification, also called speaker diarization, works out how many people are talking in a recording and which parts belong to whom. The transcript comes back as a conversation rather than a single block:

Speaker 1  So how did you first hear about the programme?
Speaker 2  A colleague mentioned it, honestly I was sceptical at first.
Speaker 1  What changed your mind?

For anyone who quotes people — researchers, journalists, lawyers, UX teams — this is the difference between a usable transcript and one you have to re-listen to.

Whisper alone does not identify speakers

This surprises many people. OpenAI's Whisper is excellent at turning speech into text, but it has no concept of who is speaking. Speaker labels require a second model, a diarization step, layered on top.

Most apps solve this by sending your audio to a server where that step runs. That is exactly what you cannot do with a confidential interview, a client call or a patient conversation. Scrieb runs both steps, transcription and speaker identification, locally on your Mac or Windows PC. Nothing leaves the device.

Where it matters most

Research interviews

Separate the interviewer from the participant, and participants from each other in focus groups. Quotes are attributable without replaying the audio.

Meetings and calls

Know who committed to what. Decisions and action items stay attached to the person who said them.

Legal and compliance

Client conversations, hearings and witness statements need attribution, and they must not be uploaded to a third party.

Journalism and podcasts

Multi-guest recordings become quotable dialogue, ready to edit or publish.

How it works in Scrieb

  1. Drop in a recording. MP3, WAV, M4A, MP4 and other common formats, in 90+ languages.
  2. Scrieb transcribes and labels speakers automatically. On your CPU or GPU, offline. Nothing to configure.
  3. Export with the labels intact. DOCX, PDF or TXT, formatted as a dialogue.

Getting the best results

Speaker identification is only as good as the audio. A few habits make a large difference, whichever tool you use:

  • Record close to the speakers. A phone in the middle of a table is better than a laptop microphone across the room.
  • Avoid crosstalk. Overlapping speech is the hardest case for any diarization model.
  • Keep the original file. Transcribe the uncompressed recording rather than a voice-memo re-export or a video with music over it.
  • Two to six speakers is the sweet spot. Large panels with many similar voices get harder to separate.

Speaker identification: local tools compared

ToolSpeaker labelsRuns offlineMac + Windows
Scrieb✓ all plans
MacWhisperPro onlyMac only
Buzz
Whisper CLITerminal
Cloud services (Otter, Turboscribe …)✗ upload requiredWeb

Full comparison: Best Local Transcription Software.

Frequently asked questions

What is speaker identification in transcription?
Speaker identification, also called speaker diarization, splits a recording by who is talking and labels each part of the transcript by speaker. Instead of one block of text, you get a conversation: Speaker 1 said this, Speaker 2 answered that.
Does Whisper identify speakers?
No. OpenAI's Whisper model only converts speech to text; it has no idea how many people are talking. Speaker labels come from a separate diarization step. Most tools run that step in the cloud. Scrieb runs both transcription and speaker identification locally on your computer.
Does speaker identification work offline?
Yes. In Scrieb the whole pipeline runs on your Mac or Windows PC. Nothing is uploaded, and after the initial model download no internet connection is needed.
How many speakers can Scrieb tell apart?
Scrieb detects the speakers in a recording automatically. It works best with recordings of two to six people where speakers do not constantly talk over each other, which covers most interviews, meetings and focus groups.
How accurate is speaker identification?
It depends mostly on the audio. Clear recordings with distinct voices and little crosstalk get very good results. Phone-quality audio, heavy background noise or several people speaking at once reduce accuracy for every tool, cloud or local.
Is speaker identification included in every plan?
Yes. Speaker identification is included in all Scrieb plans, with no minute limits. Some competitors reserve it for a Pro tier or charge per minute.
Can I export the transcript with speaker labels?
Yes. Speaker labels are kept when you export to DOCX, PDF or TXT, so the document reads like a dialogue and is ready to quote from.

Related pages

→ Interview Transcription→ Meeting Transcription→ Best Local Transcription Software→ Local Transcription→ GDPR Transcription