✳ FREE · NO UPLOAD · CHINESE

Chinese speech to text, free and private

Drop a Mandarin recording or video. Whisper runs on your device and returns an editable transcript and Chinese subtitles, 16 characters per line by default.

Public-domain LibriVox recording. Load it, then press Transcribe to try the full tool.

Whisper base, the model this tool runs, scored a 34.1% character error rate on Chinese in the Whisper paper's FLEURS test (lower is better). Expect a rough draft: many words will need fixing.

First use downloads the speech model and engine (about 97 MB, of which 72.5 MB is the model); later visits load it from your browser cache. Runs on your CPU; nothing is uploaded. Up to 100 MB and 2 hours per file; under 30 minutes is the comfortable range on a laptop.

To transcribe Mandarin Chinese for free, drop the file above; Whisper base runs in your browser and nothing is uploaded. The Whisper paper measures a 34.1% character error rate for this model on Mandarin (FLEURS), so expect a rough draft: sound-alike characters are the main errors, and the output may come out in Traditional characters.

Checked by Koldflux · updated 2026-09-25 · tested on a public-domain Chinese recording

How accurate is Chinese transcription?

Character error rate from the Whisper paper (lower is better): FLEURS, read sentences recorded under the cmn_hans_cn (Mandarin, Simplified script) locale, and Common Voice 9, volunteers with varied microphones. This tool runs Whisper base. Koldflux Studio runs Whisper large-v3-turbo, which the paper does not measure; OpenAI reports large-v3 cuts errors by 10 to 20% against large-v2, and the turbo version trades a little of that accuracy for speed.

ModelFLEURS CERCommon Voice CERVerdict (FLEURS)
Whisper base (this tool)34.1%44.9%Rough draft only
Whisper small20.8%29.4%Usable, needs a proofread
Whisper medium12.1%23.2%Good
Whisper large-v214.7%26.8%Good
Sources: Whisper paper (Radford et al., 2022), Appendix D (checked 2026-09-25) · FLEURS dataset card (Google) (checked 2026-09-25) · Whisper large-v3 model card (checked 2026-09-25) · Whisper large-v3-turbo model card (checked 2026-09-25)

What we saw on a real Chinese recording

We ran the first 30 seconds of a Mandarin reading of The Art of War (LibriVox, public domain) through the same files this page loads, with the language set to Chinese. One 30-second clip is an illustration, not a benchmark. Processing took 7.4 seconds on one CPU thread of our test server.

孫子冰法 第一張這是為 Librevox 點org 所提供的錄音一切 Librevox 的錄音都為公眾所有如果您想知道更多關於 Librevox 的信息或者提供志願服務請參看 Librevox 點org 的網站第一張史記第一孫子冰者 國之大事 死生之地 存亡之道

Simplified or Traditional, and sound-alike characters

Whisper has one Chinese setting and no switch between Simplified and Traditional; in our test it wrote Traditional characters for a Mandarin reading. If you need Simplified, convert the finished transcript with any Traditional-to-Simplified converter; the conversion does not change timestamps.

Because many characters share a pronunciation, the typical error is a correct sound with the wrong character (兵 and 冰). Accuracy is measured per character: the paper spaces out every character before scoring.

Cantonese: this model has no Cantonese setting. OpenAI added a Cantonese language token only in Whisper large-v3, so Cantonese speech here is written as if it were Mandarin-style Chinese.

Chinese subtitles and translation

Value
Characters per line (Netflix)16 characters per line (Simplified Chinese)
Reading speed (Netflix)up to 9 characters per second (adult programs)
Tool default line length16 characters
Speech to English, Whisper base (BLEU, higher is better)1.9
Speech to English, Whisper large-v2 (BLEU)18.4

Tips for Chinese

  1. Chinese to English subtitles: two stepsWhisper base scores a BLEU of 1.9 translating Chinese speech to English (large-v2: 18.4). Transcribe in Chinese, fix the characters, then translate the text.
  2. Subtitle lengthNetflix’s Simplified Chinese guide sets 16 characters per line and about 9 characters per second for adults. The tool defaults to 16 for Chinese.
  3. PunctuationThe model sometimes leaves out Chinese punctuation or uses spaces instead. Add full-width punctuation where the subtitles need it.

Limits and when to use something else

Frequently asked questions

Why is my transcript in Traditional Chinese?

Whisper chooses the script itself; there is no Simplified/Traditional option. Convert the text afterwards; timestamps are unaffected.

Does it support Cantonese?

Not properly. Whisper base has no Cantonese setting; a Cantonese token was only added in Whisper large-v3.

How accurate is it for Mandarin?

Whisper base: 34.1% character errors on FLEURS and 44.9% on Common Voice. Whisper large-v2 reaches 14.7% and 26.8%.

Can I make English subtitles from a Chinese video?

Use the two-step route: Chinese transcript here, then a text translator. Whisper base’s direct speech translation from Chinese scores a BLEU of 1.9.

WHEN YOU NEED A LARGER MODEL

Turn one idea into a week of posts

Koldflux Studio (paid, from $19/month) transcribes Mandarin uploads on our servers with Whisper large-v3-turbo (the paper measures 14.7% for the comparable large-v2), then turns the recording into Shorts and carousel posts you review, approve and export. It does not translate. It does not publish or schedule for you.

  1. 1 · Recording, idea or script
  2. 2 · Pick the angles
  3. 3 · Review and approve
  4. 4 · Export Shorts and carousels
See Koldflux Studio
Koldflux Studio: choosing which angles (Opportunities) to turn into Shorts before anything is generated
Koldflux Studio: you choose the angles before anything is made.

Language pages

Each page has the tool preset to that language, measured accuracy, our own test on a real recording and subtitle rules.