Blog · September 20, 2026 · 16 min read

Getting a Transcript Made: Your Options

Graphic title card for the article “Getting a Transcript Made: Your Options” with a stylised audio waveform with transcript.
Grafik: HumanITy

Having a transcript made means three different things, and the per-minute price is the worst yardstick for choosing between them. First the machine: automated speech recognition from 0.17 euros per audio minute at Amberscript, advertised by the vendor as "90%+ accuracy", finished in minutes. Second the human: professional transcription from 1.85 euros per minute at Amberscript and 1.99 dollars at Rev, each with a vendor promise of more than 99 percent, delivered between twelve hours and five working days. Third the middle path: machine first, human correction afterwards, from 1.75 euros per minute at Happy Scribe. The gap between roughly ten euros and roughly 110 euros per hour of audio is a factor of eleven, and it shrinks as soon as you count the rework. A service pays off where the machine does not cope with your material: poor recording, people talking over each other, strong dialect, dense technical vocabulary, or a case where the wording has to be exact. What you supply in return is the original file, a decision on the accuracy level you want, a list of proper nouns, and the legal basis that lets a third party listen to the recording at all. All prices as of September 2026.

I transcribe my own recordings with Whisper locally on my own machine. A first consultation contains numbers, names and sometimes health details. How I hand tasks to fixed roles instead of a rotating set of subscriptions is on the AI workforce page, and which data goes towards AI in my business is in AI and Privacy: What the AI Gets to See. This post is about who should use a service anyway.

The calculation that actually decides it

It isn't the per-minute price that decides, it's your own time afterwards. In their handbook Praxisbuch Interview, Transkription & Analyse (9th edition, January 2024), Dresing and Pehl put manual typing under a simple rule system at five to ten times the length of the recording as a realistic plan. The fastest person they ever measured came in at 1 to 3, but only by skipping the second correction pass. For automated speech recognition they state explicitly that there are no public studies on the time required including correction, and offer as their own observation a time saving of around 50 percent and more compared with typing it out by hand.

What comes out of that is a worked example, not a measurement, and both assumptions are on the table: five to ten times the recording length for typing it out, of which around half is saved. On that basis one hour of audio, machine plus your own correction, lands at two and a half to five hours of work. Material, rule system and quality target shift the range, and your binding number only comes from the ten-minute test at the end. On those assumptions the comparison looks like this:

Route Cost per hour of audio Your time afterwards
automated, you correct around 10 euros (0.17 euros per minute) the whole correction pass, 2.5 to 5 hours in the worked example
professional, human around 111 euros (1.85 euros per minute) re-reading the passages that matter
machine plus human proofreading around 105 euros (1.75 euros per minute) re-reading

Prices, delivery times and accuracy figures come from the providers' pricing and product pages, as of September 2026, and change there without notice.

So the difference is roughly 100 euros per hour of audio, and what you buy for it in the same example is two and a half to five hours. That works out at 20 to 40 euros per hour bought. That too is a scenario, not a market price: it inherits the assumptions above and leaves out the project and proofreading time that a service job also costs. As an order of magnitude it holds up: whether it is expensive is decided by your own hourly rate, not by a price list. For a single recording the sum is irrelevant; for ten interviews it isn't. Dresing and Pehl budget 50 to 100 working hours for ten hours of interviews under a simple rule system, which is two to four weeks at four to six hours a day.

One point applies equally to both routes: correction is always necessary. Every transcript, automated or typed by hand, contains errors after the first pass. Dresing and Pehl cite a study by Chiari from 2006 according to which untrained transcribers produce an error in almost every paragraph, and around 37 percent of those errors distort the meaning of the statement. noScribe puts the same sentence in its own documentation: review and correction are always necessary. Commissioning a service does not buy a flawless text, it buys you a much smaller remainder.

When handing it off is the right call

There are recordings where automated speech recognition works badly or not at all. Dresing and Pehl name three groups: strong dialect, Swiss German for example, several people speaking at the same time, and background noise of the kind you get in a canteen or a restaurant. On top of that, some voices defeat automated speaker separation, which you then fix by hand. Google names the same list for automatic captions on YouTube: pronunciation, accents, dialects, background noise and, explicitly, overlapping speakers.

Five cases where the human is the right answer:

  1. The recording is poor. A phone on the table, the room next door, a call recording squeezed through bandwidth compression. The machine then guesses fluently and plausibly, which makes the errors harder to spot than obvious nonsense would be.
  2. Several speakers talk over each other. Four people at a table and two of them starting at once. Speaker separation is precisely the feature that collapses first here.
  3. Strong dialect. A recording in broad Swiss German or Bavarian will not get you there with a standard model.
  4. Dense technical vocabulary. Product names, drugs, file numbers, industry jargon. Those are exactly the words that matter, and the ones the recogniser mishears.
  5. It has to be word for word. For a qualitative analysis, an expert opinion or a file that later has to show who said what, readability is not the measure, agreement with what was said is. That does not come from a human typing instead of a machine, it comes from what you agree on: the rule system (usually verbatim here), a second pass for quality assurance, a secure transfer route and, where proceedings are involved, their own requirements.

And the honest counter-case: a cleanly recorded two-person conversation in standard German that only needs to be searchable for you needs no service at all. Which tools do that job is in Transcription Software Compared, and which of them cost nothing is in Convert Audio to Text for Free.

What you have to supply

These six points have worked as a brief in my own work. They decide whether you get a transcript you can actually use:

  1. The original file, untouched. Not sent through a messenger, not converted by you, not trimmed. Every intermediate step costs audio quality, and audio quality is the whole job here. The services take video files as readily as audio.
  2. The accuracy level you want. See the next section. Without an instruction you get the provider's house style, and that is rarely the one you need.
  3. A list of proper nouns and technical terms. People, companies, products, abbreviations. Ten minutes of preparation that saves most of the later correction.
  4. The number of speakers and how to label them. Whether the transcript says "Speaker 1" or "Consultant" is your decision beforehand, not the provider's.
  5. Deadline and delivery format. Word, plain text, SRT or VTT for subtitles, JSON for further processing. The format decides your next hour of work.
  6. The legal basis and the contract for it. A basis under Article 6 GDPR, plus the conditions in Article 9 GDPR for health details and other special categories. The processing agreement must be in place before the service processes the recording, and participants receive the information on recipients or categories of recipients that fits the case. See the section after next.

The accuracy levels

Amberscript offers three variants for professional transcription, and that three-way split is common elsewhere too:

Level What happens Use for
smoothed grammar corrected, filler words removed reports, articles, minutes
clean, word for word with light cleanup word for word, small errors corrected interviews, analysis
verbatim every word, pauses and filler sounds included legal work, compliance, language analysis

Happy Scribe likewise distinguishes verbatim from clean read for its human service and includes timestamps and speaker labels. Rev offers verbatim as an add-on to the standard price.

Above those three sits the research world with its own rule systems, GAT for instance, with pitch contours and volume. Dresing and Pehl put that at around 60 times the length of the recording against 5 to 10 times for a simple system, and a GAT2 basic transcript at 18 hours per hour of material for a practised transcriber. Anyone asked for that is in a different price bracket.

A paragraph about the data

Handing a recording to a service means letting a third party process personal data. Under Article 28 GDPR that requires a data processing agreement before this processing begins, covering, among other things, processing only on documented instructions, a confidentiality obligation for the provider's staff, approval of sub-processors, and deletion or return of the data when the job ends. In the research workflow described by Dresing and Pehl, that agreement is already in place before the conversations because external disclosure is listed in the consent form. The legal basis and timing can differ in another case. Article 13 GDPR requires information on recipients or categories of recipients. And where health details or other special categories under Article 9 GDPR are involved, a basis under Article 6 alone is not enough. Anyone handling professional secrets, in a healthcare profession, in legal advice or in tax advice for example, also has section 203 of the German Criminal Code in view: subsection 3 permits disclosure to people who assist in the professional activity, as far as that is necessary, and subsection 4 makes the professional liable if they failed to ensure that the assisting person was bound to secrecy. The providers know this: Amberscript lists an NDA and a processing agreement as available and states that it stores data in Europe, Happy Scribe offers an NDA option for sensitive projects. None of this is legal advice, and when you may record and transcribe in the first place is covered in detail in Transcribing Calls: What German Law Says.

The third option: don't outsource it, hire it

Both routes above end at the same place: you have text, and more of it with every job. After thirty conversations there are thirty files nobody reads any more.

That role is filled in my business by Gustav, my AI employee for call analysis. His order of operations differs from a transcription service: speech recognition runs through Whisper locally on my own machine. After that comes a fixed anonymisation pass that turns names into roles, companies into industries, amounts into orders of magnitude and places into regions. The replacing alone is half the distance: under Art. 4(5) GDPR that is pseudonymisation, and under recital 26 such data stays personal. As a minimum check, the pass therefore asks whether role, industry, timing and a striking quote together still point back to a person. True anonymity also depends on the full context and on whether identification remains possible by means reasonably likely to be used. If in doubt, the extract stays pseudonymised and gets a purpose, a retention period and access control. Gustav's output is therefore not the transcript. He condenses many conversations into a register of patterns: the recurring questions, each with the best answer you gave to it, and out of that come a playbook, an FAQ, raw material for content and an objection library. He distils what you said yourself and does not invent advice on top. How a role like that comes into being is in Hiring an AI Employee: The Process, and which roles hold up in coaching work is in AI for Coaches: Practice Over Hype.

Frequently asked questions

What does it cost to have a transcript made?

Billing is per audio minute. Automated transcription starts at 0.17 euros per minute at Amberscript, so roughly ten euros per hour of recording. The human service starts at 1.85 euros per minute there, at 1.99 dollars at Rev and from 1.75 euros at Happy Scribe. Surcharges apply for rush delivery and extras such as anonymisation, timestamps and speaker identification. As of September 2026.

Can I get a transcript made for free?

Free gets you the machine, not the human. Whisper and noScribe run on your own computer at no cost, and the cloud services have capped free tiers for testing, ten minutes of AI transcription at Happy Scribe for example. The correction still sits with you. Which routes genuinely cost nothing and where their limits are is in Convert Audio to Text for Free.

Is AI enough, or do I need a human?

For a cleanly recorded conversation in standard German with clear turn-taking, the machine is enough for your own use. A human pays off with poor recordings, people speaking simultaneously, strong dialect, dense technical vocabulary, and whenever the transcript has to match the words exactly. The test is not the advertised accuracy rate, it's the time you spend afterwards.

How do I get a transcript of a video or a YouTube video?

Transcription services accept video files directly, so you don't have to extract the audio first. On YouTube the platform generates automatic captions in more than 100 languages, German included, which can serve as a text base. Google itself recommends always reviewing automatic captions and editing the parts that were not transcribed properly.

How long does a transcription job take?

Automated transcription delivers in minutes. For the human service Amberscript states, depending on the page, three to five working days as standard and 24 hours or one working day as a rush job, while Rev promises twelve hours or less. Budget the time for your own read-through as well, it comes up either way.

How to take this further

Take ten minutes of a typical recording of your own, not a sample file, and have it transcribed automatically once. Then time how long you need until you would show that text to somebody. That single number answers the question better than any price list: multiply it by the volume you produce in a year and compare it with 1.85 euros a minute. If what comes out of that is that the bottleneck isn't the conversion but the working-through afterwards, Gustav is available as a finished package in my community, with the anonymisation pass, consent templates and an invented demo transcript to practise on. The templates are samples without warranty and do not replace a lawyer's review of your case. Everything about it is at Community.

Kevin Welter

Kevin Welter

Developer, IT architect, author of technical books (Kubernetes, cloud infrastructures) and speaker. Runs his business with an AI workforce of fourteen AI employees and shows solo business owners in his community how to hire their first AI employee.

More about AI employees

Your first AI employee up and running within an hour

Join the community