Picking foreign-language content used to mean guessing. A book is labeled "B1," a news site is "for intermediate learners," and neither tells you the thing you actually need to know: whether you, specifically, with your specific vocabulary, can follow it without drowning. Comprehension research has a real answer to that question, and LingoBlend now surfaces it directly instead of leaving you to estimate.
The research: what "readable" actually means#
Two thresholds show up repeatedly in vocabulary and reading research, both stated as a percentage of the running words on a page you already know — not the size of your vocabulary in general, but coverage of this specific text. Laufer's work on the 95% threshold found that comprehension becomes workable with some contextual guessing once you recognize about 19 words in every 20. Hu and Nation (2000) pushed further and found that comfortable, dictionary-free reading of fiction generally needs closer to 98% — about one unknown word per short paragraph rather than one per line.
| Coverage of running words | Reading experience | Source |
|---|---|---|
| ~98% | Comfortable, unassisted — reading for pleasure | Hu & Nation, 2000 |
| ~95% | Workable with effort — assisted comprehension, some guessing | Laufer's 95% threshold work |
| ~90% | Frustrating — heavy dictionary use, slow going | — |
| Below ~90% | A wall — more decoding than reading | — |
The gap between 95% and 98% looks small on paper and is enormous in practice: it's the difference between one unfamiliar word every twenty and one every fifty, which is the difference between a text you fight through and one you actually enjoy. A CEFR label can't tell you which side of that line you're on for a given piece of writing, because the label describes the text's general difficulty, not your personal coverage of its specific vocabulary.
Picking the right text: the "% known" badge#
For content that already exists — graded stories and the chapters of a downloaded classic book — LingoBlend now shows you the coverage number directly rather than a difficulty label you have to translate into a guess. A badge like "94% known" is the same measurement the research above is built on: the share of that text's running content words that are ones you've actually proven you know.
Two details make the number mean what it says rather than flatter you:
- It counts occurrences, not just distinct words. A text where you know 90% of the distinct vocabulary but the 10% you don't know happens to repeat constantly would still read as difficult — so the badge weights by how often each word actually appears, exactly the way the 95%/98% research thresholds are defined.
- Function words are excluded and inflections are resolved. The count runs over content words only — grammar glue like articles and pronouns is stripped out per-language, so the number reflects vocabulary knowledge, not sentence structure. And a story using despierta counts toward your mastery of despertar the same way it does everywhere else in the app, via the vocabulary graph's form index — you're not penalized for a text choosing a conjugated form over the dictionary one.
Like the vocabulary size stat, the badge only appears once you have enough proven review history for it to mean something, and only on texts long enough that a percentage isn't noise from a handful of lucky or unlucky words. Below either bar, nothing shows — a badge on a paragraph-long story would be measuring almost nothing.
Making any text the right difficulty: the suggested blend percentage#
Graded stories and books are pre-written, so a badge is the right tool: it scores content that already exists. But most of what people actually want to read — an article, a newsletter, a recipe — isn't graded for anyone, and you can't badge your way to the right difficulty on content nobody has pre-scored. That's the other half of the same problem, and LingoBlend solves it the opposite way: instead of measuring a fixed text against you, it adjusts the text to you.
When you paste text into Blend, the app tokenizes it on-device and checks it against your mastered vocabulary — the same underlying mastery data the coverage badge reads. From how much of your pasted text you already know, it estimates the blend percentage that would land new-word density in the productive range — the comprehensible-input sweet spot researchers describe as roughly a 5–15% band of new words, comfortably inside what your existing vocabulary can anchor. The result appears as a small tappable chip, something like "Try 35%," next to the percentage slider.
That suggestion is exactly that — a suggestion. It never moves the slider for you. You tap it if you want it, adjust further if you don't, and it's just as easy to ignore entirely and set your own number the way you always could. If you're a newer learner without enough review history yet for a text-specific estimate, the suggestion falls back to a flat number based on your self-reported level from level-aware blending — still a starting point, never a mandate. If neither signal exists yet, you simply see no suggestion, and the slider works exactly as it always has.
Two sides of the same idea#
Both features are answering the same question — "how hard is this for me?" — from opposite directions. The badge scores a fixed text so you can pick one that fits. The blend slider, the mechanism behind the diglot weave method LingoBlend automates, adjusts any text until it fits, which is the genuinely unique move: you're not limited to a library someone else graded, because you're not reading text at a fixed difficulty in the first place.
That's also why blended texts never carry a coverage badge of their own. A badge answers "how hard is this text as written" — but a blended text isn't fixed; its difficulty is the percentage you chose. The slider already tells you the answer a badge would have to compute.