Conversational fluency in Japanese takes roughly 3,000 words, enough for around 95% coverage of everyday speech, while reading unedited books and newspapers comfortably points toward 10,000 words or more alongside close to the full jōyō set of 2,136 kanji. Japanese is the language where the standard question breaks down most, because vocabulary and writing-system knowledge are separate axes and neither one alone predicts whether you can read a page.
Japanese needs two counts, not one#
In Spanish or German, knowing a word and being able to read it are nearly the same skill. In Japanese they come apart. You can know the spoken word kansatsu perfectly and stall on 観察 in print, and you can know a kanji's meaning and readings and still fail to parse a compound built from it. Vocabulary is how many distinct words you know, spoken or written. Kanji is how many of the roughly 2,000 characters in general use you can read in context. Progress on one supports the other but does not substitute for it, so any honest answer has to give both figures.
What counts as a word in Japanese#
Japanese text has no spaces, so word boundaries are not given by the writing system. They are imposed by a segmentation standard, and standards disagree. The Balanced Corpus of Contemporary Written Japanese, the basis for Tono, Yamazaki and Maekawa's A Frequency Dictionary of Japanese (2013), defines both short unit words and long unit words, and the same text yields different totals under each. That ambiguity has no equivalent in European languages.
| Unit | Definition | Japanese example |
|---|---|---|
| Token | Every running word under a chosen segmentation | 食べる, 食べる, 食べた = 3 tokens |
| Type | Each distinct form, counted once | 食べる, 食べた = 2 types |
| Lemma | A headword plus its inflected forms | 食べる covers 食べます, 食べた, 食べない, 食べられる |
| Word family | A headword plus derivations | 食べる plus 食べ物, 食事 through shared kanji |
The word-family concept, which does so much work in English and Romance vocabulary research, fits Japanese awkwardly. Derivation happens largely through kanji compounding: 電 (electric) plus 車 (vehicle) gives 電車 (train), and 電話 (telephone) comes nearly free once the character is known. Character knowledge partly substitutes for word-by-word learning in a way that has no European parallel. Katakana loanwords form another semi-free band for English speakers, though meanings drift, so マンション is an apartment and not a mansion.
The coverage thresholds behind "fluent"#
Vocabulary research measures coverage, the share of running words in a text that a reader already knows. 95% coverage is described as the floor for adequate comprehension when context can fill the gaps (Laufer, 1989), about one unknown word in twenty. 98% coverage is the level Hu and Nation (2000) associated with independent reading, about one unknown word in fifty. Nation's 2006 analysis put the vocabulary needed for 98% at roughly 8,000-9,000 word families in English, with less required for speech than for writing.
Japanese figures are usually quoted higher, around 10,000 or more for comfortable reading, for two reasons. The family unit collapses, so counts are given in words rather than families and inflate accordingly, and written Japanese leans on Sino-Japanese compounds that are individually rare but collectively constant. Treat the 10,000 figure as the same threshold in a looser unit, not as evidence that Japanese demands more learning than French.
| Most frequent words | Approx. text coverage | What it typically unlocks |
|---|---|---|
| 1,000 | ~75-80% | Survival Japanese, the gist of simple conversation |
| 2,000 | ~85-90% | Everyday exchanges with gaps; graded readers |
| 3,000 | ~93-95% | Comfortable spoken fluency on familiar topics |
| 5,000 | ~96% | Manga, drama with subtitles, simplified news |
| 10,000+ | ~98% | Unedited novels and newspapers, read for pleasure |
The kanji axis: jōyō, kyōiku, and coverage#
The Japanese government's jōyō kanji list, revised in 2010, contains 2,136 characters and defines the set expected of an educated adult reader. A subset of 1,006 kyōiku kanji is taught across the six years of elementary school. Newspapers and general publishing broadly respect the jōyō boundary, and characters outside it usually carry furigana.
Character frequency counts of newspaper and general-purpose corpora show the same Zipfian lopsidedness that Zipf described in 1949. Approximate figures, and they are approximate:
| Kanji known | Approx. character coverage | Practical effect |
|---|---|---|
| 100 | ~40-50% | Signs, dates, numbers, basic labels |
| 500 | ~80% | Simple graded text, much manga with furigana |
| 1,000 | ~95% | Most sentences readable with a few gaps |
| 2,136 (full jōyō) | ~99% | Newspapers and general publishing |
Note what that 95% at 1,000 kanji does and does not mean. Character coverage is not word coverage. One unrecognized character in a two-character compound usually blocks the whole word, so the effective failure rate at the word level is higher than the character figure suggests. This is why the last thousand kanji still matter even though they look statistically marginal.
Japanese levels: JLPT and CEFR#
The JLPT stopped publishing official vocabulary and kanji lists after its 2010 redesign. The figures below come from the pre-2010 test content specifications and from widely circulated third-party analyses, so they describe the shape of the exam rather than an official syllabus. The Japan Foundation's JF Standard maps Japanese teaching onto the CEFR, which is where the level column comes from.
| JLPT | Approx. CEFR | Approx. vocabulary | Approx. kanji |
|---|---|---|---|
| N5 | A1 | ~800 | ~100 |
| N4 | A2 | ~1,500 | ~300 |
| N3 | A2-B1 | ~3,500-4,000 | ~650 |
| N2 | B1-B2 | ~6,000 | ~1,000 |
| N1 | B2-C1 | ~10,000 | ~2,000 |
N1 is not native-level, despite its reputation. It certifies reading and listening only, sets no speaking or writing bar, and passing it still leaves plenty of novels difficult.
Why every number here is an estimate#
The unit is contested, the corpora differ, and the exam figures are unofficial. Frequency lists drawn from subtitle collections such as OpenSubtitles skew conversational and undercount the Sino-Japanese compounds that dominate written registers, while newspaper corpora do the reverse. Knowing a word is not binary either, and Japanese adds a split: you may know a word by ear, by sight, or both. Finally, coverage is not comprehension. Keigo, ellipsis, and unstated subjects can leave you knowing every word in a sentence and still missing who did what to whom.
How to actually build both counts#
Lists alone are slow on the vocabulary axis and merely tedious on the kanji axis. Two mechanisms carry the load.
Volume of comprehensible reading. Krashen's Input Hypothesis (1985) holds that acquisition follows from understanding messages slightly above your level. In Japanese this does double duty, since every page of reading also drills character recognition in context, which no isolated kanji drill reproduces. Material with furigana lowers the entry barrier considerably, which is the practical argument in reading manga to learn Japanese.
Scheduled review. Reviewing at expanding intervals counters the forgetting curve Ebbinghaus documented in 1885, a finding confirmed across dozens of experiments reviewed by Cepeda and colleagues in 2006. Kanji in particular decay fast without it. The practical guide to spaced repetition covers the scheduling.
LingoBlend puts both in one loop. Paste any text, set the blend percentage, and Japanese words are woven into material you already understand, which is Burling's 1968 diglot weave applied with an AI engine. Tapping a blended word gives its meaning, grammar, and base form, and saved words feed an SM-2 spaced-repetition queue across five review games with pronunciation audio, which matters more here than in a language where spelling predicts the reading. The Japanese learning guide covers where to start, and the Spanish vocabulary targets show how differently this plays out in a European language.