✳ FREE · NO UPLOAD · FRENCH

French speech to text, free in your browser

Drop a French recording or video. Whisper runs on your device and returns an editable, timestamped transcript and French subtitles.

Public-domain LibriVox recording. Load it, then press Transcribe to try the full tool.

Whisper base, the model this tool runs, scored a 28.5% word error rate on French in the Whisper paper's FLEURS test (lower is better). Expect a usable draft that needs a careful proofread.

First use downloads the speech model and engine (about 97 MB, of which 72.5 MB is the model); later visits load it from your browser cache. Runs on your CPU; nothing is uploaded. Up to 100 MB and 2 hours per file; under 30 minutes is the comfortable range on a laptop.

To transcribe French audio for free, drop the file above; the Whisper base model runs in your browser and nothing is uploaded. The Whisper paper measures a 28.5% word error rate for this model on French (FLEURS), so treat the result as a usable draft: most sentences are right, but homophones and names need a proofread.

Checked by Koldflux · updated 2026-09-25 · tested on a public-domain French recording

How accurate is French transcription?

Word error rate from the Whisper paper (lower is better): FLEURS, read sentences recorded under the fr_fr (France) locale, and Common Voice 9, volunteers with varied microphones. This tool runs Whisper base. Koldflux Studio runs Whisper large-v3-turbo, which the paper does not measure; OpenAI reports large-v3 cuts errors by 10 to 20% against large-v2, and the turbo version trades a little of that accuracy for speed.

ModelFLEURS WERCommon Voice WERVerdict (FLEURS)
Whisper base (this tool)28.5%37.3%Usable, needs a proofread
Whisper small15%22.7%Good
Whisper medium8.7%16%Good
Whisper large-v28.3%13.9%Good
Sources: Whisper paper (Radford et al., 2022), Appendix D (checked 2026-09-25) · FLEURS dataset card (Google) (checked 2026-09-25) · Whisper large-v3 model card (checked 2026-09-25) · Whisper large-v3-turbo model card (checked 2026-09-25)

What we saw on a real French recording

We ran the first 30 seconds of the LibriVox recording of La Fontaine’s Fables, book one (LibriVox, public domain) through the same files this page loads, with the language set to French. One 30-second clip is an illustration, not a benchmark. Processing took 6.6 seconds on one CPU thread of our test server.

Ceci est un enregistrement libre-y-vox. Tous nos enregistrements appartiennent tout de même public. Pour vous renseigner à notre sujet ou pour participer, rendez-vous sur libre-y-vox.org … Fable de la Fontaine, livre premier, de Jean de la Fontaine, à mon Seigneur le Dauphin.

Where French transcripts go wrong

French has many words that sound the same and are spelled differently: a and à, ces and ses, verb endings in -é, -er and -ez. A speech model hears one sound and has to guess the spelling from context, so these are the errors to look for first.

Liaisons and elisions can also move word boundaries, as in our test where “au domaine public” turned into “tout de même public”. The paper’s French number was measured on speakers from France (FLEURS fr_fr); Canadian, Belgian and African accents are handled by the same setting.

French subtitles and translation

Value
Characters per line (Netflix)42 characters per line
Reading speed (Netflix)up to 17 characters per second (adult programs)
Tool default line length42 characters
Speech to English, Whisper base (BLEU, higher is better)13.1
Speech to English, Whisper large-v2 (BLEU)32.2

Tips for French

  1. Proofread homophones, not every wordScan for a/à, ou/où, ces/ses and past participles; they are the bulk of the errors in a clean French recording.
  2. French to English: transcribe firstWhisper can translate French speech straight to English, but this small model does it roughly (BLEU 13.1 against 32.2 for the largest model). A French transcript translated as text reads much better.
  3. Subtitle speed for FrenchNetflix’s French guide sets 42 characters per line and about 17 characters per second for adults. Keep long cues on two balanced lines.

Limits and when to use something else

Frequently asked questions

How accurate is free French transcription?

For the model this tool runs, the Whisper paper reports 28.5% word errors on FLEURS and 37.3% on Common Voice. The largest model in the paper, large-v2, reaches 8.3% and 13.9%. So this is a draft you edit, not a finished transcript.

Can I translate French audio to English here?

Yes, choose “Translate the speech to English”. The result is a rough translation into English only. For a publishable result, export the French transcript and translate the text.

Does it add punctuation and capitals?

Yes, Whisper writes sentences with punctuation and capital letters. It does not add speaker names or labels.

Is my recording uploaded?

No. The audio is decoded and transcribed inside your browser tab; only the model files are downloaded, once.

FROM TRANSCRIPT TO POSTS

Turn one idea into a week of posts

Koldflux Studio (paid, from $19/month) transcribes French uploads on our servers with Whisper large-v3-turbo (the paper measures 8.3% for the comparable large-v2), then turns the recording into Shorts and carousel posts you review, approve and export. It does not publish or schedule for you.

  1. 1 · Recording, idea or script
  2. 2 · Pick the angles
  3. 3 · Review and approve
  4. 4 · Export Shorts and carousels
See Koldflux Studio
Koldflux Studio: choosing which angles (Opportunities) to turn into Shorts before anything is generated
Koldflux Studio: you choose the angles before anything is made.

Language pages

Each page has the tool preset to that language, measured accuracy, our own test on a real recording and subtitle rules.