It is the first question everybody asks and the one that gets the worst answers. "Three thousand words for fluency." "You only need a thousand." "Native speakers know twenty thousand." All three are quoted constantly, none of them is wrong exactly, and none of them helps, because they are answers to different questions.
The useful version of the question is: how much of what I hear will I understand if I know the N commonest words? That has an actual shape, it is the same shape in every language, and once you have seen it you cannot unsee it in your own learning.
The curve
Rank every word in a language by how often it is spoken. Now walk down that list adding up what share of all speech you have covered. The curve rises viciously fast and then flattens almost to horizontal.
The rough shape, for informal spoken language:
| Words known | Share of speech you have met before |
|---|---|
| 100 | around half |
| 300 | around 70% |
| 1,000 | around 85% |
| 2,000 | around 90% |
| 5,000 | around 95% |
| 10,000+ | around 98% |
The figures vary by language, by register, and by how you count a "word" — whether go, goes, going and went are one item or four makes a large difference, and the research convention is to count word families rather than forms. The numbers most often quoted come from Paul Nation's vocabulary work and the corpus studies around it. But the shape is not in dispute, and the shape is the point.
Two consequences follow, and they pull in opposite directions.
Consequence one: the first three hundred words are absurdly valuable
Half of everything anyone says to you is carried by a hundred or so words. Not interesting words — the, and, is, have, that, not, what, do, the pronouns, the numbers, the handful of verbs that do all the work. They are boring, they are hard to make flashcards for because their meaning is structural rather than pictorial, and they are worth more than the next two thousand put together.
This is why the vocabulary you pick up incidentally is so badly distributed. Left to yourself, you look up the words that stand out — and words stand out precisely because they are rare. You notice the word for "windowsill" and you never look up the word that means "anyway", because you half-recognise it and it does not feel like a word that needs learning. So you accumulate a deck of interesting, uncommon words while the words carrying most of the language stay at sixty per cent known, which in listening is the same as not known.
Any serious approach has to push against that instinct deliberately. Ours does it by counting: we rank words by how often they are actually said across every transcript in our catalogue, and the review questions come in that order rather than in the order a learner happened to tap. A word you miss enters your deck by itself. The instinct is still there — you will still tap "windowsill" — but the system is not built on it.
Consequence two: the last few per cent take years
The curve flattens, and the flat part is very long. Getting from 95% to 98% coverage means roughly doubling your vocabulary, from about five thousand words to ten thousand and beyond. That is the difference between a few months and several years.
And it matters more than it looks. At 95% coverage you are meeting one unknown word in every twenty — about one per sentence, which is enough to derail a joke, a punchline or a critical clause. The research on reading suggests you need somewhere around 98% coverage to read comfortably without constant lookups; listening is harsher still, because you cannot go back.
So "you only need a thousand words" and "you need ten thousand words" are both true. A thousand gets you the gist of ordinary speech. Ten thousand gets you the meaning, including the bits that were the point.
What this means for how you spend your time
Front-load the common words, aggressively. The first few hundred are so much more valuable than anything else that it is worth doing them explicitly rather than waiting to meet them. This is the one phase where drilling beats immersion on pure efficiency.
Then stop drilling and start listening. Past the first thousand or so, the bottleneck is no longer which words you have met but how fast you can retrieve them, and that comes from meeting them in context, repeatedly, at speed. A deck of five thousand cards is a five-thousand-item chore that still leaves you unable to follow a conversation.
Judge material by coverage, not by topic. If you are meeting more than about one unknown word in ten, the piece is above you and you will get less from it than from something easier. This is a real measurement you can take on any transcript: count the words you do not know, divide by the total. Over 10% and it is a study text, not listening practice.
Stop counting after a while. Word counts are motivating early and misleading later, because they stop tracking the thing you care about. Knowing a word is not binary — it runs from "I have seen it" through "I can recognise it when reading" to "I hear it at speed in a sentence and never notice I understood it". Only the last of those is worth anything in conversation, and no counter measures it.
Why we count our own words instead of using a list
Published frequency lists exist and they are good work. We do not use them, for three reasons worth stating because they change the numbers.
Register. Most frequency lists are built from written corpora — books, newspapers, subtitles. The vocabulary of a novel is not the vocabulary of someone talking to a camera about their lunch. If we teach short-form video, the useful ranking is the one taken from short-form video.
Licensing. The good published lists mostly carry share-alike terms. A corpus built from transcripts we already hold carries none.
Honesty. Counting our own catalogue makes coverage a measured number rather than a claim. The rank at which cumulative coverage crosses 80% is literally that language's eighty-twenty line, in the material we actually serve, and it moves as the catalogue grows.
That last point cuts both ways and it is worth being straight about the weakness. A small catalogue gives a noisy tail: the first few hundred ranks settle quickly — Zipf's law is reliable at the head — but rank four thousand moves every time a video is published. So we act on bands rather than on ranks. "Inside the half of all speech carried by the fewest words" is a decision worth making. "Rank 412" is not a fact yet.
The short answer
If you want one number: the commonest three hundred words of a language carry roughly seventy per cent of ordinary conversation, and you can learn them in a few weeks. That is the single best trade available in language learning, and almost nobody takes it deliberately.
Everything after that is a long, flat climb, and the way up it is hours of understanding real language — not more cards. If you want the argument for that, it is in comprehensible input and how to get some; if you want the reason your vocabulary is not turning into listening ability, it is in why you can read the sentence and still not catch a word.