LingoBlend

How to Build Vocabulary by Reading (Without Losing the Plot)

Reading builds vocabulary slowly on its own. Here is what the coverage and repeated-encounter research actually says, and how to fix the two things that make it slow.

MethodNikola Artukov9 min readUpdated

"Just read more" is the most common vocabulary advice in language learning, and it is correct in the way that "just eat less" is correct — true, insufficient, and unhelpful about the part people actually get stuck on. This article covers what reading does and does not do to your vocabulary, the two variables that decide which one you get, and what to change when it is not working.

What reading actually does to vocabulary#

The mechanism is called incidental vocabulary acquisition: you meet an unknown word inside a sentence, infer something about it from context, and a trace of it stays. You were not trying to learn it. You were trying to find out what happened next.

The research on this is unusually consistent about two things. First, it works — readers do reliably pick up words they were never taught. Second, it works slowly per word. Studies of incidental learning from context generally find only a modest chance of acquiring any particular unknown word from a single encounter, and what you get from that encounter is partial: a rough sense of meaning, often without the spelling, the gender, the register, or the ability to produce it.

That sounds like a case against reading. It is not, and the reason is arithmetic. A 60,000-word novel puts you through tens of thousands of word encounters. Even at a low rate per encounter, and even counting only unknown words, the total is large — and none of it required you to make a card, choose a word, or maintain a queue. No deliberate method produces that volume, because no one will sit through it.

The right way to hold both facts at once: reading is a poor way to learn a specific word and an excellent way to learn thousands of unspecified ones. Deliberate study is the reverse. This is why the two are complements rather than competitors, and why the strongest setups run both.

Variable one: coverage#

Coverage is the fraction of the running words on a page that you already know. It is the single best predictor of whether reading will teach you anything, and it behaves less like a dial than a cliff.

The rough thresholds from Paul Nation's work and the studies around it:

CoverageWhat it feels likeWhat happens to learning
~80%One unknown word every fiveIncomprehensible. No acquisition.
~90%One in tenHeavy decoding, constant lookup, exhausting
~95%One in twentyReadable with support; comprehension mostly holds
~98%One in fiftyComfortable unassisted reading; guessing from context works

The reason the drop is so sharp is that inferring a word from context requires the context to be intact. At 98%, an unknown word sits in a sentence you fully understand, and the sentence does most of the work. At 90%, the surrounding words are also uncertain, so there is nothing solid to infer from — you are guessing from a guess. This is why "read harder material to learn faster" fails: past a certain difficulty you are not learning more slowly, you are learning close to nothing while working much harder.

Reaching 98% coverage on unsimplified text takes roughly 8,000–9,000 word families, which is the number behind most "how many words do I need" answers. Around 3,000 families cover about 95% of everyday text.

Variable two: repeated encounters#

The second variable is how many times a word comes back before you forget it.

Estimates in the literature vary with what you count as "knowing" a word, but the general finding is that a single meeting is rarely enough, and that substantial, durable knowledge tends to require something on the order of a dozen encounters, spread out rather than clustered. Ebbinghaus established the shape of the forgetting curve in 1885: without repetition, retention drops steeply and early.

Natural text distributes those encounters brutally unevenly. High-frequency words come back constantly, which is why they get learned almost for free. Mid-frequency words — precisely the band that gates real reading — might appear twice in a whole novel, hundreds of pages apart. The forgetting curve wins that race every time.

This is the actual gap in "just read more". Reading supplies encounter one. It does not reliably supply encounters two through twelve for the words that matter most, and no amount of additional reading fixes the distribution problem, because the distribution is a property of the language.

Fixing coverage: four options#

If you are below ~95% on what you want to read, you have four moves.

Read something easier. Graded readers, children's books, news written for learners. Effective and often unsatisfying — the price is that controlled material is rarely material you would have chosen.

Read something you already know. A translated novel you have read in your own language, or an article whose subject you know well. Prior knowledge of the content substitutes for vocabulary knowledge, and it is the cheapest coverage boost available.

Narrow reading. Stay inside one author, one series, or one topic. Vocabulary repeats hard within a domain, so coverage climbs quickly and the repeated-encounter problem partly solves itself. Reading five articles about the same news story beats five articles about five stories.

Change the text's difficulty directly. Rather than finding a text at your level, adjust one you want to read. Pre-glossing the least common words is one form. Partial translation is another: LingoBlend takes text in a language you already know and swaps a percentage of the words — you set it with a slider — into the language you are learning, so coverage is a setting rather than a property of the book. That inverts the usual constraint: instead of hunting for material at 95%, you take material you want and put it there. It is the diglot weave technique, described by Robbins Burling in 1968, covered in what is the diglot weave method.

Fixing repetition: capture and schedule#

The second fix is smaller and less negotiable. Words the text will not repeat enough need to come back on a schedule.

Capture at the moment of curiosity, not later. A word you resolve in the sentence that made you want it arrives with meaning, register and grammar attached. The same word transcribed into a list that evening arrives naked, and you have already lost why it mattered. The capture has to cost roughly one second or you will stop doing it — a tap in a reader, a right-click on a web page, a share from your phone.

Save far less than you meet. The instinct is to save everything unknown, and it is the reliable way to end up with 400 saved words and 40 known ones. A working filter: save it if it appeared twice, or if the sentence collapsed without it, or if you can imagine using it. Otherwise let it go — it will come back if it matters.

Put the saved words on a spaced schedule. This is the part that supplies encounters two through twelve. Any SM-2-style scheduler does the job; the algorithm matters far less than whether you open it. Ten words a day, reviewed, beats fifty saved and abandoned. See spaced repetition for language learning for what a schedule needs to do, and flashcards that actually work for what to put on the card.

Keep the context. A saved word with its original sentence is a substantially better card than a word pair, because it preserves the collocation and the register that made the meaning clear.

Putting it together#

A version that survives ordinary weeks:

  1. Choose material by interest first, then check coverage. Interest is what gets you to chapter nine. Read two pages and count your stops: more than about five unknown words per hundred and the text is too hard, so either pick something easier or lower its difficulty directly.
  2. Read in blocks with a shape. A chapter, an article, a section. Something that ends.
  3. Read for the plot, not for the words. Look up only what recurs or what breaks the sentence. The rest is allowed to stay fuzzy — that fuzziness is what incidental acquisition is made of.
  4. Save selectively, in the moment. One tap, and only for words that passed the filter.
  5. Review the next day, briefly. Ten minutes. This is where the reading converts into vocabulary.
  6. Raise the difficulty when it gets comfortable. Coverage climbs as you go, so material that was too hard three months ago is now in range.

The whole thing is one loop: reading supplies the first encounter and the context, review supplies the rest of the encounters, and the growing vocabulary raises coverage, which makes the next text easier and the next first encounters more productive. It compounds, which is why people who read consistently pull away from people who study harder.

Related reading: intensive vs extensive reading for how close and wide reading divide the work, comprehensible input explained for the theory underneath coverage, how many words to be fluent in Spanish for what the targets look like in one language, and best vocabulary builder apps if you want the tooling comparison.

Frequently asked questions

Can you learn vocabulary just by reading?

Yes, and it is how most of a large vocabulary gets built, in a first language as well as a second. The catch is that acquisition per encounter is low and partial, so it needs volume to work. Reading alone will get you there eventually; reading plus a small amount of scheduled review on the words you deliberately save gets you there considerably faster, because it supplies the repeat encounters natural text distributes too thinly.

How many times do I need to see a word to remember it?

Estimates vary with what counts as knowing a word, but a single encounter is rarely enough and durable knowledge tends to need something on the order of a dozen spaced encounters. High-frequency words get those for free from ordinary reading. Mid-frequency words, which are the ones actually blocking you, often appear twice in an entire book — which is the gap a review schedule exists to close.

Should I look up every unknown word while reading?

No. Looking up everything turns reading into decoding, destroys the pace, and removes the context-inference that makes incidental acquisition work in the first place. A workable rule is to look a word up if it appears twice or if the sentence stops making sense without it. If you are hitting that rule constantly, the text is too hard rather than your discipline being too weak.

What percentage of words do I need to know to read a book?

About 95% of running words for comprehension with support, and roughly 98% for comfortable unassisted reading, per Nation's coverage research. Below about 90% comprehension collapses and very little is learned. In practice, 98% means roughly one unknown word every fifty — around one or two per page.

Is reading better than flashcards for vocabulary?

They do different jobs and the comparison is a false choice. Flashcards are efficient at making a specific word stick and supply no context or volume. Reading supplies enormous volume and rich context but cannot guarantee any particular word returns before you forget it. Used together — read for encounters, review what you saved — each covers the other's weakness.

Share this articleXLinkedInRedditEmail

Nikola Artukov

Builder of LingoBlend. Writes about reading as a way into a language — the methods, the research behind them, and the practical workflows that make them fit into an ordinary week.

More about the author

Related reading

Start learning a new language today

Join LingoBlend and turn any text into a personalized language lesson. Free to start, no credit card required.