German voice to text, free and private
Drop a German recording or video. The speech model runs on your device and gives you an editable transcript and subtitles.
Public-domain LibriVox recording. Load it, then press Transcribe to try the full tool.
Whisper base, the model this tool runs, scored a 17.9% word error rate on German in the Whisper paper's FLEURS test (lower is better). Expect a usable draft that needs a careful proofread.
First use downloads the speech model and engine (about 97 MB, of which 72.5 MB is the model); later visits load it from your browser cache. Runs on your CPU; nothing is uploaded. Up to 100 MB and 2 hours per file; under 30 minutes is the comfortable range on a laptop.
To turn German speech into text for free, drop the file above; Whisper base transcribes it inside your browser without uploading it. The Whisper paper measures a 17.9% word error rate for this model on German (FLEURS), so expect a usable draft in which long compound words and names need checking.
Checked by Koldflux · updated 2026-09-25 · tested on a public-domain German recording
How accurate is German transcription?
Word error rate from the Whisper paper (lower is better): FLEURS, read sentences recorded under the de_de (Germany) locale, and Common Voice 9, volunteers with varied microphones. This tool runs Whisper base. Koldflux Studio runs Whisper large-v3-turbo, which the paper does not measure; OpenAI reports large-v3 cuts errors by 10 to 20% against large-v2, and the turbo version trades a little of that accuracy for speed.
| Model | FLEURS WER | Common Voice WER | Verdict (FLEURS) |
|---|---|---|---|
| Whisper base (this tool) | 17.9% | 24.5% | Usable, needs a proofread |
| Whisper small | 10.2% | 13% | Good |
| Whisper medium | 6.5% | 8.5% | Good |
| Whisper large-v2 | 4.5% | 6.4% | Good |
What we saw on a real German recording
We ran the first 30 seconds of Schiller’s poem “Phantasie an Laura” (LibriVox, public domain) through the same files this page loads, with the language set to German. One 30-second clip is an illustration, not a benchmark. Processing took 5.9 seconds on one CPU thread of our test server.
Fantasie an Laura von Friedrich Schiller, gelesen für LibriVox.org Meine Laura nenne mir den Wirbel, der an Körperkörpermächte greist, nenne meine Laura mir den Zauber, der zum Geist gewaltigt zwingt den Geist.- Detected as German with 99.1% confidence.
- Schiller’s line “der an Körper Körper mächtig reißt” became “der an Körperkörpermächte greist”: unusual word order and old poetic German are where the model guesses.
- The English translation of this clip first got stuck repeating “the spirit of the spirit…”; the tool’s fallback pass fixed the loop automatically.
German compounds and capitals
Whisper capitalises German nouns and usually joins compounds correctly in everyday speech. Rare or invented compounds are the weak spot: the model may split them, join the wrong parts, or, as in our poem test, glue several words into one.
The paper’s German number comes from speakers in Germany reading prepared sentences (FLEURS de_de). Swiss German and strong dialects are far from that test, so expect clearly worse results on them.
German subtitles and translation
| Value | |
|---|---|
| Characters per line (Netflix) | 42 characters per line |
| Reading speed (Netflix) | up to 17 characters per second (adult programs) |
| Tool default line length | 42 characters |
| Speech to English, Whisper base (BLEU, higher is better) | 13.7 |
| Speech to English, Whisper large-v2 (BLEU) | 34.6 |
Tips for German
- Search for your compoundsProduct names and technical compounds are the words most likely to be split or merged. Check them first.
- Use the fallback noticeIf the tool says it re-decoded part of the audio, proofread that part: it means the first pass started repeating itself.
- German subtitlesNetflix’s German guide allows 42 characters per line and about 17 characters per second for adults. Long German words make balanced two-line cues harder; shorten the line length setting if a word overflows.
Limits and when to use something else
- Dialects, overlapping voices and music under the speech push errors well above the 17.9% benchmark.
- English translation is available but rough with this model (BLEU 13.7 against 34.6 for large-v2).
- Speed depends on your device. A laptop handles a 30-minute file comfortably; phones may run out of memory on long files. Hard limits: 100 MB and 2 hours per file.
- No speaker labels: Whisper writes one stream of text. It also does not mark music, laughter or background sounds reliably.
Frequently asked questions
How good is free German transcription?
The paper measures 17.9% word errors for Whisper base on German FLEURS and 24.5% on Common Voice; the large-v2 model reaches 4.5% and 6.4%. In practice: most sentences right, some words to fix.
Does it work with Swiss German or Austrian German?
It uses the same German setting for all of them. The benchmark is standard German from Germany, so dialect-heavy speech will have noticeably more errors.
Can it make German subtitles for a video?
Yes. Drop the video, transcribe, fix the text, and download SRT or VTT. The files work on YouTube and in the subtitle tools linked below.
Does it recognise who is speaking?
No. Whisper produces one stream of text with timestamps and no speaker labels.
Turn one idea into a week of posts
Koldflux Studio (paid, from $19/month) transcribes German uploads on our servers with Whisper large-v3-turbo (the paper measures 4.5% for the comparable large-v2), then turns the recording into Shorts and carousel posts you review, approve and export. It does not publish or schedule for you.
- 1 · Recording, idea or script
- 2 · Pick the angles
- 3 · Review and approve
- 4 · Export Shorts and carousels

Language pages
Each page has the tool preset to that language, measured accuracy, our own test on a real recording and subtitle rules.