Japanese audio to text, free in your browser
Drop a Japanese recording or video. Whisper runs on your device and writes an editable Japanese transcript and subtitles, 13 characters per line by default.
Public-domain LibriVox recording. Load it, then press Transcribe to try the full tool.
Whisper base, the model this tool runs, scored a 22.8% character error rate on Japanese in the Whisper paper's FLEURS test (lower is better). Expect a usable draft that needs a careful proofread.
First use downloads the speech model and engine (about 97 MB, of which 72.5 MB is the model); later visits load it from your browser cache. Runs on your CPU; nothing is uploaded. Up to 100 MB and 2 hours per file; under 30 minutes is the comfortable range on a laptop.
To transcribe Japanese for free, drop the file above; Whisper base runs in your browser and nothing is uploaded. The Whisper paper measures a 22.8% character error rate for this model on Japanese (FLEURS): a usable draft, with kanji homophones as the main thing to fix. Its English translation of Japanese is poor, so translate the text instead.
Checked by Koldflux · updated 2026-09-25 · tested on a public-domain Japanese recording
How accurate is Japanese transcription?
Character error rate from the Whisper paper (lower is better): FLEURS, read sentences recorded under the ja_jp (Japan) locale, and Common Voice 9, volunteers with varied microphones. This tool runs Whisper base. Koldflux Studio runs Whisper large-v3-turbo, which the paper does not measure; OpenAI reports large-v3 cuts errors by 10 to 20% against large-v2, and the turbo version trades a little of that accuracy for speed.
| Model | FLEURS CER | Common Voice CER | Verdict (FLEURS) |
|---|---|---|---|
| Whisper base (this tool) | 22.8% | 24.2% | Usable, needs a proofread |
| Whisper small | 12% | 14% | Good |
| Whisper medium | 7.1% | 10.5% | Good |
| Whisper large-v2 | 5.3% | 9.1% | Good |
What we saw on a real Japanese recording
We ran the first 30 seconds of the LibriVox recording of Sōseki’s “Meian” (LibriVox, public domain) through the same files this page loads, with the language set to Japanese. One 30-second clip is an illustration, not a benchmark. Processing took 6.5 seconds on one CPU thread of our test server.
名案、夏名祖席こちらはリブリボクスです。リブリボクスの残はすべてパブリックドメインです。ボランケアについてなど詳しくはサイトをご覧ください。URL、リブリボクス、ドットオーグ。- Detected as Japanese with 98.4% confidence.
- The title 明暗 (meian) came out as 名案, a different word with the same reading; the author 夏目漱石 became 夏名祖席. Kanji homophones are the typical error.
- The English translation began “This is the rib-leaf box”: Whisper base’s Japanese-to-English translation is not usable.
Kanji, kana and spaces
Japanese has no spaces between words, so accuracy is measured per character: the paper spaces out every character before scoring, which makes its Japanese “WER” a character error rate. The output has no spaces either, and the tool wraps subtitles by character.
Most errors are a right sound written with the wrong kanji, because many Japanese words share a reading. Katakana loanwords and names can also drift (“LibriVox” became リブリボクス). Read the transcript for meaning, not just for sound.
Japanese subtitles and translation
| Value | |
|---|---|
| Characters per line (Netflix) | 13 full-width characters per line (horizontal) |
| Reading speed (Netflix) | up to 4 characters per second |
| Tool default line length | 13 characters |
| Speech to English, Whisper base (BLEU, higher is better) | 1.5 |
| Speech to English, Whisper large-v2 (BLEU) | 18.9 |
Tips for Japanese
- Japanese to English subtitles: two stepsWhisper base scores a BLEU of 1.5 translating Japanese speech to English (large-v2 reaches 18.9). Transcribe in Japanese here, fix the kanji, then translate the text with a text translator and keep the timestamps from the SRT.
- Use Japanese subtitle limitsNetflix’s Japanese guide allows 13 full-width characters per line for horizontal subtitles and about 4 characters per second. The tool defaults to 13 for Japanese; the generic subtitle checker’s 20 characters-per-second rule does not apply.
- Fix names onceNames are written with many possible kanji. Correct the first occurrence, then use your editor’s find-and-replace on the downloaded file.
Limits and when to use something else
- Keigo-heavy business speech and dialects such as Kansai-ben are further from the benchmark than read speech.
- Do not rely on the built-in English translation for Japanese; it is the weakest part of this small model.
- Speed depends on your device. A laptop handles a 30-minute file comfortably; phones may run out of memory on long files. Hard limits: 100 MB and 2 hours per file.
- No speaker labels: Whisper writes one stream of text. It also does not mark music, laughter or background sounds reliably.
Frequently asked questions
Can I get English subtitles for a Japanese video here?
Technically yes, but with this small model they are not usable: the paper measures a BLEU of 1.5 for Japanese-to-English. Transcribe in Japanese instead, correct it, and translate the text.
How accurate is Japanese transcription?
Whisper base: 22.8% character errors on FLEURS and 24.2% on Common Voice. The largest model in the paper, large-v2, reaches 5.3% and 9.1%.
Why are there no spaces in the transcript?
Japanese is written without spaces between words, and Whisper writes it that way. Subtitles are wrapped by character count instead of by word.
Does it write furigana or romaji?
No. The output is ordinary Japanese text in kanji and kana.
Turn one idea into a week of posts
Koldflux Studio (paid, from $19/month) transcribes Japanese uploads on our servers with Whisper large-v3-turbo (the paper measures 5.3% for the comparable large-v2), then turns the recording into Shorts and carousel posts you review, approve and export. It does not translate. It does not publish or schedule for you.
- 1 · Recording, idea or script
- 2 · Pick the angles
- 3 · Review and approve
- 4 · Export Shorts and carousels

Language pages
Each page has the tool preset to that language, measured accuracy, our own test on a real recording and subtitle rules.