The three options
Human transcription service
You upload, a person types. $1–$2.50 per audio minute, days of turnaround, best accuracy on bad audio, true verbatim available.
AI transcription service
You upload, a model transcribes on their servers. $0.10–$0.25 per minute or a capped monthly plan. Fast. The audio still leaves your hands.
Local transcription software
A model runs on your own computer. Flat subscription or free open-source tools. Fast on a GPU, and the audio never leaves the machine.
The maths
One-hour interviews, typical rates, standard turnaround.
| Interviews | Human service | AI service | Scrieb (local) |
|---|---|---|---|
| 1 | $60–$150 | $6–$15 | $24.99 for the month, unlimited |
| 10 | $600–$1,500 | $60–$150 | $24.99 |
| 25 | $1,500–$3,750 | $150–$375 | $24.99–$49.98 over one or two months |
| Ongoing, 10 hours a month | $7,200–$18,000 a year | $720–$1,800 a year | $169.99 a year |
The service costs scale with every recording. The software cost doesn't. Past a handful of interviews, local software is cheaper than the cheapest AI service, and the gap to human services is an order of magnitude.
What money doesn't show
- Confidentiality. Both kinds of service receive your audio. For a consent form that promises no third-party access, or for GDPR and HIPAA, that is the whole question. Local software has no third party.
- Turnaround. A human service returns in days. Local transcription of a one-hour file takes minutes on a GPU and runs while you do something else.
- Accuracy. Human beats AI on bad audio and true verbatim. On a clean recording the difference is a read-through, which you do anyway before coding.
- Speaker labels. Human services include them. AI services usually do. Local software: Scrieb yes, most free tools no.
- Control. With a service the transcript exists in someone else's account until they delete it. With software it exists where you put it.
A simple rule
- Sensitive recordings, or a consent form that mentions third parties: local software, no discussion.
- More than five interviews: local software on cost alone.
- Terrible audio you cannot re-record, or certified verbatim required: human service.
- One short, non-sensitive file and no suitable computer: AI service.
If local is the answer, the local tools comparison covers the free options as well as Scrieb.
Pricing
Flat price, no minute limits. The 25-interview dissertation costs one or two months.
Frequently asked questions
- How much does interview transcription cost per minute?
- Human transcription services typically charge between $1 and $2.50 per audio minute, so $60 to $150 for a one-hour interview, more for rush turnaround, poor audio or verbatim style. AI-based services charge roughly $0.10 to $0.25 per minute or a monthly plan with a minute cap. Local software such as Scrieb is a flat subscription with no per-minute charge.
- Is a human transcription service more accurate?
- For difficult audio, heavy accents or true verbatim, yes: a good human transcriber still beats any model. For clear recordings the gap is small, and AI transcription is accurate enough for thematic analysis with a read-through. Many human services now run AI first and have a person correct it.
- Which option is confidential?
- Any service, human or AI, receives your audio and is a third party under your consent form and under GDPR or HIPAA. Local software keeps the audio on your own computer, so no third party is involved. That is the deciding factor for clinical, legal and sensitive research recordings.
- How fast is each option?
- Human services: one to five days, faster at a premium. AI services: minutes to an hour after upload. Local software: minutes per hour of audio on a GPU, roughly real time on a laptop without one, running unattended.
- Can local software handle speaker labels and timestamps?
- Scrieb labels speakers and keeps timecodes locally. Free local tools vary: Buzz has no speaker labels, WhisperX has them but runs in a terminal.
- When is a service still the right choice?
- When you need certified or true-verbatim transcripts, when the audio is poor and you cannot re-record, or when you have no computer capable of running a model. For most interview studies, none of those apply.