To be conversationally fluent in German you need roughly 3,000 high-frequency word families, which covers about 95% of the words in ordinary speech and writing. Reading novels and newspapers without a dictionary points to around 8,000-9,000 families (Nation, 2006). German complicates this more than most languages, because its habit of welding nouns together means the question "how many words" has no stable answer until you decide what a word is.
What counts as a word in German#
The same German paragraph produces four different totals depending on the unit you count.
| Unit | Definition | German example |
|---|---|---|
| Token | Every running word, counted each time | spricht, spricht, sprechen = 3 tokens |
| Type | Each distinct written form, counted once | spricht, sprechen = 2 types |
| Lemma | A headword plus its inflected forms | sprechen covers spreche, sprichst, sprach, gesprochen |
| Word family | A lemma plus transparent derivations | sprechen plus Sprache, Sprecher, Gespräch, aussprechen |
Two features of German make the family unit especially useful. Separable-prefix verbs split across a clause, so anrufen surfaces as ich rufe dich an, and a beginner reading that sentence sees a verb and a stray particle rather than one lemma. Strong verb ablaut changes the stem itself, so singen, sang, gesungen look like three items and are one. When someone quotes a German vocabulary target, they almost always mean lemmas or families. Jones and Tschirner's A Frequency Dictionary of German (2006) is built on roughly the top 4,000 lemmas, and its authors report that those few thousand headwords account for the large majority of running words in the corpus.
Compound nouns break the count#
German forms compounds productively, which means speakers manufacture new nouns on the spot by stacking existing ones. Handschuh is a hand-shoe, a glove. Geschwindigkeitsbegrenzung is a speed limit. Krankenversicherungskarte is a health insurance card. None of these needs to be learned as a separate item if you already know the parts, because the meaning is compositional and usually obvious.
This wrecks naive vocabulary statistics in two directions at once. A frequency list counts every compound as its own type, so German corpora show far more distinct types than English corpora of the same size, with an enormous tail of compounds appearing exactly once. That inflates any "German has X words" claim. At the same time, the learning burden is smaller than the type count suggests, because a compound you have never seen is often instantly transparent. The practical consequence: judge your German vocabulary by families and stems, not by dictionary entries, and treat unfamiliar long nouns as a decoding exercise rather than a memorization task. A learner who knows Kranken-, Versicherung, and Karte has already earned the third word for free.
Compounds are not always transparent, and that is worth flagging honestly. Handschuh is guessable; Wahrzeichen (landmark) and Augenblick (moment) are idiomatic and have to be learned. But the transparent cases dominate, especially in technical and administrative German.
The coverage thresholds behind "fluent"#
Vocabulary research measures coverage, the share of running words in a text that a reader already knows. Two thresholds recur. 95% coverage is often described as the floor for adequate comprehension when context can carry the remainder (Laufer, 1989), roughly one unknown word in twenty. 98% coverage is the level Hu and Nation (2000) associated with genuinely independent reading, roughly one unknown word in fifty. Nation's 2006 analysis put the vocabulary needed for 98% at about 8,000-9,000 word families for written text and rather less, around 6,000-7,000, for spoken language, because speech recycles a smaller core.
How much text the top 1,000, 2,000, and 5,000 words cover#
Frequency is brutally uneven. Zipf (1949) described the shape: a few items appear constantly, most appear almost never.
| Most frequent word families | Approx. text coverage | What it typically unlocks |
|---|---|---|
| 1,000 | ~80-85% | Survival German, the gist of simple conversation |
| 2,000 | ~90% | Everyday exchanges with gaps; graded readers |
| 3,000 | ~95% | Comfortable spoken fluency; most daily topics |
| 5,000 | ~96-97% | Simplified news, subtitled film with context |
| 8,000-9,000 | ~98% | Unedited novels and newspapers, read for pleasure |
The top band in German is function words and workhorse verbs: der, die, und, sein, haben, werden, können, müssen, machen, gehen. Learning them is unavoidable and enormously profitable. The reverse is the sobering part: the climb from 95% to 98% means acquiring several thousand progressively rarer words that you will meet only occasionally. That is a property of the frequency distribution, not a failure of your study method. The same curve shows up in every language studied, which is why the Spanish vocabulary targets sit in nearly the same place.
German vocabulary by CEFR level#
The CEFR describes what a learner can do and deliberately specifies no vocabulary sizes. German has an unusually good sanity check, though, because the Goethe-Institut publishes wordlists for its A1, A2, and B1 exams, and the B1 list runs to roughly 2,400 items. That maps closely to the estimates below, which draw on vocabulary-size research (Milton, 2009).
| CEFR level | Approx. word families | What it typically supports |
|---|---|---|
| A1 | 500-1,000 | Introductions, numbers, shopping, simple personal facts |
| A2 | 1,000-2,000 | Routine exchanges, past tense narration, short simple texts |
| B1 | 2,000-3,000 | Travel and workplace situations, clear standard speech, graded readers |
| B2 | 4,000-5,000 | Abstract discussion, press reading with occasional lookups |
| C1 | 6,000-8,000 | Professional and academic use, most literature |
| C2 | 9,000-10,000+ | Near-native range, including idiom, register, and specialist prose |
Why every number here is an estimate#
Treat all of the above as ranges. The unit dominates: counting families rather than types can change a German total by a factor of several, thanks to compounds. The corpus matters next, since a list derived from a subtitle collection such as OpenSubtitles weights conversational vocabulary very differently from one built on Die Zeit. Knowing a word is not binary, so passive vocabulary always exceeds active vocabulary. And coverage is not comprehension: you can know 98% of the words in a German insurance policy and still not know what you agreed to.
How to actually build a German vocabulary that size#
Lists alone are slow and forgettable. Two evidence-backed mechanisms do the work when combined.
Volume of comprehensible reading. Krashen's Input Hypothesis (1985) holds that acquisition happens when you understand messages slightly beyond your current level. Reading German you can mostly follow puts high-frequency families in front of you repeatedly in natural context, and it is also where compound-decoding becomes automatic. The distinction between careful study and high-volume reading is covered in intensive versus extensive reading.
Scheduled review. Reviewing at expanding intervals counters the forgetting curve Ebbinghaus published in 1885 far more efficiently than rereading, a finding confirmed across dozens of experiments reviewed by Cepeda and colleagues in 2006. The practical case for spaced repetition walks through the scheduling.
LingoBlend combines the two. Paste a German article or chapter, set the blend percentage, and target-language words are woven into text you already understand, which is Burling's 1968 diglot weave with a modern engine. Tapping a blended word shows its meaning, its grammar, and its base form, so gesprochen resolves to sprechen and separable verbs stop looking like two unrelated pieces. Saved words then flow into an SM-2 spaced-repetition queue feeding five review games. The German learning guide is a sensible starting point.