"Just read more" is the most common vocabulary advice in language learning, and it is correct in the way that "just eat less" is correct — true, insufficient, and unhelpful about the part people actually get stuck on. This article covers what reading does and does not do to your vocabulary, the two variables that decide which one you get, and what to change when it is not working.
What reading actually does to vocabulary#
The mechanism is called incidental vocabulary acquisition: you meet an unknown word inside a sentence, infer something about it from context, and a trace of it stays. You were not trying to learn it. You were trying to find out what happened next.
The research on this is unusually consistent about two things. First, it works — readers do reliably pick up words they were never taught. Second, it works slowly per word. Studies of incidental learning from context generally find only a modest chance of acquiring any particular unknown word from a single encounter, and what you get from that encounter is partial: a rough sense of meaning, often without the spelling, the gender, the register, or the ability to produce it.
That sounds like a case against reading. It is not, and the reason is arithmetic. A 60,000-word novel puts you through tens of thousands of word encounters. Even at a low rate per encounter, and even counting only unknown words, the total is large — and none of it required you to make a card, choose a word, or maintain a queue. No deliberate method produces that volume, because no one will sit through it.
The right way to hold both facts at once: reading is a poor way to learn a specific word and an excellent way to learn thousands of unspecified ones. Deliberate study is the reverse. This is why the two are complements rather than competitors, and why the strongest setups run both.
Variable one: coverage#
Coverage is the fraction of the running words on a page that you already know. It is the single best predictor of whether reading will teach you anything, and it behaves less like a dial than a cliff.
The rough thresholds from Paul Nation's work and the studies around it:
| Coverage | What it feels like | What happens to learning |
|---|---|---|
| ~80% | One unknown word every five | Incomprehensible. No acquisition. |
| ~90% | One in ten | Heavy decoding, constant lookup, exhausting |
| ~95% | One in twenty | Readable with support; comprehension mostly holds |
| ~98% | One in fifty | Comfortable unassisted reading; guessing from context works |
The reason the drop is so sharp is that inferring a word from context requires the context to be intact. At 98%, an unknown word sits in a sentence you fully understand, and the sentence does most of the work. At 90%, the surrounding words are also uncertain, so there is nothing solid to infer from — you are guessing from a guess. This is why "read harder material to learn faster" fails: past a certain difficulty you are not learning more slowly, you are learning close to nothing while working much harder.
Reaching 98% coverage on unsimplified text takes roughly 8,000–9,000 word families, which is the number behind most "how many words do I need" answers. Around 3,000 families cover about 95% of everyday text.
Variable two: repeated encounters#
The second variable is how many times a word comes back before you forget it.
Estimates in the literature vary with what you count as "knowing" a word, but the general finding is that a single meeting is rarely enough, and that substantial, durable knowledge tends to require something on the order of a dozen encounters, spread out rather than clustered. Ebbinghaus established the shape of the forgetting curve in 1885: without repetition, retention drops steeply and early.
Natural text distributes those encounters brutally unevenly. High-frequency words come back constantly, which is why they get learned almost for free. Mid-frequency words — precisely the band that gates real reading — might appear twice in a whole novel, hundreds of pages apart. The forgetting curve wins that race every time.
This is the actual gap in "just read more". Reading supplies encounter one. It does not reliably supply encounters two through twelve for the words that matter most, and no amount of additional reading fixes the distribution problem, because the distribution is a property of the language.
Fixing coverage: four options#
If you are below ~95% on what you want to read, you have four moves.
Read something easier. Graded readers, children's books, news written for learners. Effective and often unsatisfying — the price is that controlled material is rarely material you would have chosen.
Read something you already know. A translated novel you have read in your own language, or an article whose subject you know well. Prior knowledge of the content substitutes for vocabulary knowledge, and it is the cheapest coverage boost available.
Narrow reading. Stay inside one author, one series, or one topic. Vocabulary repeats hard within a domain, so coverage climbs quickly and the repeated-encounter problem partly solves itself. Reading five articles about the same news story beats five articles about five stories.
Change the text's difficulty directly. Rather than finding a text at your level, adjust one you want to read. Pre-glossing the least common words is one form. Partial translation is another: LingoBlend takes text in a language you already know and swaps a percentage of the words — you set it with a slider — into the language you are learning, so coverage is a setting rather than a property of the book. That inverts the usual constraint: instead of hunting for material at 95%, you take material you want and put it there. It is the diglot weave technique, described by Robbins Burling in 1968, covered in what is the diglot weave method.
Fixing repetition: capture and schedule#
The second fix is smaller and less negotiable. Words the text will not repeat enough need to come back on a schedule.
Capture at the moment of curiosity, not later. A word you resolve in the sentence that made you want it arrives with meaning, register and grammar attached. The same word transcribed into a list that evening arrives naked, and you have already lost why it mattered. The capture has to cost roughly one second or you will stop doing it — a tap in a reader, a right-click on a web page, a share from your phone.
Save far less than you meet. The instinct is to save everything unknown, and it is the reliable way to end up with 400 saved words and 40 known ones. A working filter: save it if it appeared twice, or if the sentence collapsed without it, or if you can imagine using it. Otherwise let it go — it will come back if it matters.
Put the saved words on a spaced schedule. This is the part that supplies encounters two through twelve. Any SM-2-style scheduler does the job; the algorithm matters far less than whether you open it. Ten words a day, reviewed, beats fifty saved and abandoned. See spaced repetition for language learning for what a schedule needs to do, and flashcards that actually work for what to put on the card.
Keep the context. A saved word with its original sentence is a substantially better card than a word pair, because it preserves the collocation and the register that made the meaning clear.
Putting it together#
A version that survives ordinary weeks:
- Choose material by interest first, then check coverage. Interest is what gets you to chapter nine. Read two pages and count your stops: more than about five unknown words per hundred and the text is too hard, so either pick something easier or lower its difficulty directly.
- Read in blocks with a shape. A chapter, an article, a section. Something that ends.
- Read for the plot, not for the words. Look up only what recurs or what breaks the sentence. The rest is allowed to stay fuzzy — that fuzziness is what incidental acquisition is made of.
- Save selectively, in the moment. One tap, and only for words that passed the filter.
- Review the next day, briefly. Ten minutes. This is where the reading converts into vocabulary.
- Raise the difficulty when it gets comfortable. Coverage climbs as you go, so material that was too hard three months ago is now in range.
The whole thing is one loop: reading supplies the first encounter and the context, review supplies the rest of the encounters, and the growing vocabulary raises coverage, which makes the next text easier and the next first encounters more productive. It compounds, which is why people who read consistently pull away from people who study harder.
Related reading: intensive vs extensive reading for how close and wide reading divide the work, comprehensible input explained for the theory underneath coverage, how many words to be fluent in Spanish for what the targets look like in one language, and best vocabulary builder apps if you want the tooling comparison.