You measure vocabulary size by testing a random sample of words drawn from each frequency band of a language, then scaling your score up to the size of the band. Sample ten words from the third 1,000-word band, get eight right, and the test credits you with roughly 800 words there. The arithmetic is simple. The hard part sits underneath it: deciding what counts as a word.
What you are actually counting#
Three units are in circulation, and they produce wildly different totals from identical knowledge. A surface form is every distinct written form, so run, runs, running, ran counts as four. A lemma groups a headword with its inflections in one part of speech, so those four collapse into one while runner stays separate. A word family folds in transparent derivations too, making run, runs, running, ran, runner, rerun a single family. Bauer and Nation (1993) set out graded levels of affixation so researchers could state which derivations they counted.
| Unit | run, runs, running, ran, runner, rerun counts as | Typical use |
|---|---|---|
| Surface form | 6 | Corpus token counts, spell-checkers |
| Lemma | 2 (run, runner) | Dictionaries, most bilingual word lists |
| Word family | 1 | Nation's coverage research, size tests |
This is why "I know 5,000 words" is nearly unfalsifiable on its own. One learner is honestly describable as knowing 3,000 word families, around 5,000 lemmas, or well over 10,000 surface forms. The gap widens in morphologically rich languages, where a single Finnish, Turkish, or Russian noun appears in dozens of forms, so surface-form counts flatter those learners. Before comparing your number to a published target, check that both sides use the same unit. Most serious targets, including those in how many words to be fluent in Spanish, are stated in word families.
How vocabulary size tests work#
A frequency list is cut into bands of 1,000 word families, ordered from most to least common. A size test draws a small fixed sample from each band, usually a multiple-choice item asking for the meaning of the target word in a short neutral context. Your score in a band estimates the proportion of that band you know, and the estimates are summed.
The method is sound but the sample is thin. Ten items per band means each item carries a hundred words of weight, so one lucky guess moves the estimate by 100. Band-level results are noisy while the total is far more stable. Treat any result as a range of a few hundred words and re-measure months apart.
Checking for overclaiming with non-words#
The cheapest format is a yes/no checklist: you see words and tick the ones you know. Uncorrected, it is badly inflated, because learners tick words that merely feel familiar and nothing penalizes optimism. The standard fix is to seed the list with plausible-looking non-words built from the language's real spelling. Meara and Buxton (1987) established the approach: a tick on an invented word measures overclaiming directly, so the score becomes real-word hits minus the false-alarm rate.
LexTALE (Lemhöfer and Broersma, 2012) is the best-known modern version, with English, Dutch, and German editions. It reports a proficiency percentage rather than a size in words, which is the more defensible claim.
The standard tests: VST and VLT#
Two instruments from Paul Nation's group dominate the field, and they answer different questions. The Vocabulary Size Test (Nation and Beglar, 2007) is the totals test: 140 multiple-choice items, ten from each of the first fourteen 1,000-family bands, each correct answer worth 100 families, so the raw score is multiplied by 100. Bilingual versions exist in many first languages, removing the confound of parsing an English definition to prove you know an English word.
The Vocabulary Levels Test (Nation, 1983, revised by Schmitt, Schmitt and Clapham, 2001) is the diagnostic. It probes specific bands, traditionally 2,000, 3,000, and 5,000 plus academic and 10,000 sections, by matching words to short definitions. Its output is a profile: solid at 2,000, patchy at 3,000, thin at 5,000. That tells you where to spend the next three months, which a total never does.
Why self-assessment overestimates#
Recognition is much easier than recall, and both are easier than use. Nation's framing separates a word's form, meaning, and use; self-rating captures only the first. You see contemplar, feel a flicker of recognition, and tick it, without being able to produce it in a sentence or separate it from a near neighbor. Cognates amplify this: a reader coming from English recognizes hundreds of Romance words on sight without knowing which are false friends.
Multiple choice adds its own inflation, since four options return 25% on blind guessing and elimination pushes that higher. Serious estimates therefore correct for guessing, add non-word controls, or use a recall format. For a private check, cover the translations on fifty of your saved words and write the meanings out; the gap against your recognition score is your overclaiming rate.
Coverage: the measure that actually predicts reading#
If your goal is reading, the useful quantity is not your total but the share of running words on a page you already know. Laufer (1989) identified roughly 95% coverage as the threshold where comprehension becomes workable with a dictionary. Hu and Nation (2000) found comfortable unassisted comprehension of fiction generally requires closer to 98%. Nation (2006) estimated that 98% of written text takes 8,000 to 9,000 word families, while spoken text needs roughly 6,000 to 7,000, since speech draws on a narrower range.
Those percentages are easier to judge as densities. At 95% one word in twenty is unknown, about one per line of a paperback. At 98% it is one in fifty, roughly one per short paragraph, which context usually resolves without breaking your flow. That is the difference between studying a text and reading it. The same logic underpins comprehensible input and explains why intermediate readers stall, a pattern covered in the intermediate plateau.
What a given vocabulary size lets you read#
The figures below are approximate, drawn from Nation's coverage research on written English. Other languages shift the boundaries and individual texts can sit outside them.
| Word families | Approx. coverage of written text | What that realistically supports |
|---|---|---|
| 1,000 | ~75–80% | Graded readers at A1, labels, simple dialogue |
| 2,000 | ~85–90% | A2 graded readers, familiar everyday topics |
| 3,000 | ~90% | Most everyday conversation, B1 adapted texts |
| 4,000–5,000 | ~95% | Ordinary news and non-fiction with a dictionary |
| 8,000–9,000 | ~98% | Unmodified novels and journalism, read for pleasure |
| ~15,000–20,000 | Beyond 98% | The range usually estimated for educated native readers |
Native-speaker estimates vary across studies because they too depend on the counting unit, so read that last row as an order of magnitude. Mapping bands onto proficiency labels is imprecise but useful, and CEFR levels explained sets out how the descriptors line up.
Measuring coverage on your own texts#
You can skip formal testing and measure the thing you care about directly. Take a page you want to read, count off 100 running words, and mark the ones you cannot understand. Two or fewer puts you near the 98% zone, so just read. Around five puts you near 95%: readable with support and productive for vocabulary. More than ten and you are decoding rather than reading, which makes the text the problem, not you.
That count also tells you how to configure a tool. LingoBlend sets coverage deliberately: you paste a text you already wanted to read and a slider controls what share of the words appear in your target language, so unknown-word density stays in the productive range rather than wherever the author left it. Tapped words save with their grammar context and return on a spaced-repetition schedule. The rest of LingoBlend's reading and review tooling is on the features page.