✳ FREE · NO UPLOAD · JAPANESE

Japanese audio to text, free in your browser

Drop a Japanese recording or video. Whisper runs on your device and writes an editable Japanese transcript and subtitles, 13 characters per line by default.

Public-domain LibriVox recording. Load it, then press Transcribe to try the full tool.

Whisper base, the model this tool runs, scored a 22.8% character error rate on Japanese in the Whisper paper's FLEURS test (lower is better). Expect a usable draft that needs a careful proofread.

First use downloads the speech model and engine (about 97 MB, of which 72.5 MB is the model); later visits load it from your browser cache. Runs on your CPU; nothing is uploaded. Up to 100 MB and 2 hours per file; under 30 minutes is the comfortable range on a laptop.

To transcribe Japanese for free, drop the file above; Whisper base runs in your browser and nothing is uploaded. The Whisper paper measures a 22.8% character error rate for this model on Japanese (FLEURS): a usable draft, with kanji homophones as the main thing to fix. Its English translation of Japanese is poor, so translate the text instead.

Checked by Koldflux · updated 2026-09-25 · tested on a public-domain Japanese recording

How accurate is Japanese transcription?

Character error rate from the Whisper paper (lower is better): FLEURS, read sentences recorded under the ja_jp (Japan) locale, and Common Voice 9, volunteers with varied microphones. This tool runs Whisper base. Koldflux Studio runs Whisper large-v3-turbo, which the paper does not measure; OpenAI reports large-v3 cuts errors by 10 to 20% against large-v2, and the turbo version trades a little of that accuracy for speed.

ModelFLEURS CERCommon Voice CERVerdict (FLEURS)
Whisper base (this tool)22.8%24.2%Usable, needs a proofread
Whisper small12%14%Good
Whisper medium7.1%10.5%Good
Whisper large-v25.3%9.1%Good
Sources: Whisper paper (Radford et al., 2022), Appendix D (checked 2026-09-25) · FLEURS dataset card (Google) (checked 2026-09-25) · Whisper large-v3 model card (checked 2026-09-25) · Whisper large-v3-turbo model card (checked 2026-09-25)

What we saw on a real Japanese recording

We ran the first 30 seconds of the LibriVox recording of Sōseki’s “Meian” (LibriVox, public domain) through the same files this page loads, with the language set to Japanese. One 30-second clip is an illustration, not a benchmark. Processing took 6.5 seconds on one CPU thread of our test server.

名案、夏名祖席こちらはリブリボクスです。リブリボクスの残はすべてパブリックドメインです。ボランケアについてなど詳しくはサイトをご覧ください。URL、リブリボクス、ドットオーグ。

Kanji, kana and spaces

Japanese has no spaces between words, so accuracy is measured per character: the paper spaces out every character before scoring, which makes its Japanese “WER” a character error rate. The output has no spaces either, and the tool wraps subtitles by character.

Most errors are a right sound written with the wrong kanji, because many Japanese words share a reading. Katakana loanwords and names can also drift (“LibriVox” became リブリボクス). Read the transcript for meaning, not just for sound.

Japanese subtitles and translation

Value
Characters per line (Netflix)13 full-width characters per line (horizontal)
Reading speed (Netflix)up to 4 characters per second
Tool default line length13 characters
Speech to English, Whisper base (BLEU, higher is better)1.5
Speech to English, Whisper large-v2 (BLEU)18.9

Tips for Japanese

  1. Japanese to English subtitles: two stepsWhisper base scores a BLEU of 1.5 translating Japanese speech to English (large-v2 reaches 18.9). Transcribe in Japanese here, fix the kanji, then translate the text with a text translator and keep the timestamps from the SRT.
  2. Use Japanese subtitle limitsNetflix’s Japanese guide allows 13 full-width characters per line for horizontal subtitles and about 4 characters per second. The tool defaults to 13 for Japanese; the generic subtitle checker’s 20 characters-per-second rule does not apply.
  3. Fix names onceNames are written with many possible kanji. Correct the first occurrence, then use your editor’s find-and-replace on the downloaded file.

Limits and when to use something else

Frequently asked questions

Can I get English subtitles for a Japanese video here?

Technically yes, but with this small model they are not usable: the paper measures a BLEU of 1.5 for Japanese-to-English. Transcribe in Japanese instead, correct it, and translate the text.

How accurate is Japanese transcription?

Whisper base: 22.8% character errors on FLEURS and 24.2% on Common Voice. The largest model in the paper, large-v2, reaches 5.3% and 9.1%.

Why are there no spaces in the transcript?

Japanese is written without spaces between words, and Whisper writes it that way. Subtitles are wrapped by character count instead of by word.

Does it write furigana or romaji?

No. The output is ordinary Japanese text in kanji and kana.

FROM TRANSCRIPT TO POSTS

Turn one idea into a week of posts

Koldflux Studio (paid, from $19/month) transcribes Japanese uploads on our servers with Whisper large-v3-turbo (the paper measures 5.3% for the comparable large-v2), then turns the recording into Shorts and carousel posts you review, approve and export. It does not translate. It does not publish or schedule for you.

  1. 1 · Recording, idea or script
  2. 2 · Pick the angles
  3. 3 · Review and approve
  4. 4 · Export Shorts and carousels
See Koldflux Studio
Koldflux Studio: choosing which angles (Opportunities) to turn into Shorts before anything is generated
Koldflux Studio: you choose the angles before anything is made.

Language pages

Each page has the tool preset to that language, measured accuracy, our own test on a real recording and subtitle rules.