There is a specific and very deflating moment that arrives about a year into a language. You can read a newspaper paragraph with a dictionary. You know several thousand words. You can construct a sentence with a relative clause in it and get the agreement right. Then somebody in a video says one short sentence at ordinary speed and you catch the first word and nothing else.
This is not a vocabulary problem and it is not a talent problem. It is a rate problem, and it is worth understanding precisely, because the wrong diagnosis leads to months of the wrong work.
The number your textbook audio is hiding
Speech has a measurable rate. For languages that write spaces between words, the natural unit is words per second; for Japanese or Chinese, which do not, you have to count characters instead, because "word" is not a thing the writing system marks.
The rates that matter are not the ones you have been listening to. Course audio is recorded by voice actors who have been asked to be intelligible, and it typically runs somewhere near the bottom of the natural range. Ordinary conversation between two native speakers who are not thinking about it runs much faster — in most languages close to double.
These are the reference rates we use internally to grade clips, in the unit each language is counted in:
| Language | Deliberate delivery | Ordinary conversation |
|---|---|---|
| Spanish | 2.2 words/sec | 3.6 words/sec |
| French | 2.1 words/sec | 3.4 words/sec |
| German | 1.9 words/sec | 3.0 words/sec |
| Korean | 1.3 words/sec | 2.3 words/sec |
| Japanese | 3.2 chars/sec | 5.6 chars/sec |
Two things to take from that table.
The first is the ratio. In every one of those languages, ordinary conversation is roughly 1.6 to 1.8 times the rate of careful, deliberate speech. If your listening practice has all been at the deliberate end, you have been training on a signal that is most of a second per sentence slower than the one you are failing at.
The second is that the numbers are not comparable across languages. Spanish at 3.6 words a second and German at 3.0 are not telling you that Spanish speakers talk faster in any meaningful sense — German packs far more into each word. A single threshold across languages would mean "gentle" in one and "impossible" in another, which is exactly why anything that claims to grade content has to hold a reference per language rather than one number for everyone. Ours are informed estimates rather than laboratory measurements, and we say so; what matters for grading is the ratio between them, which is stable.
What speed actually does to comprehension
Here is why an extra word per second is so much worse than it sounds.
Understanding speech in a second language is not one process. It is at least three, running at once:
- Segmentation. Finding where one word stops and the next starts. In your first language this is automatic and you are not aware it is happening. In a second language it is work, and speech contains almost none of the gaps you imagine it does — the spaces you "hear" between words are mostly reconstructed by a brain that already knows the words.
- Retrieval. Getting from a sound to a meaning. For a word you know well this is instant. For a word you learned from a flashcard three weeks ago it might take a full second, and a full second is three or four words of speech.
- Assembly. Holding the pieces in working memory long enough to build the sentence.
The third one is where it collapses. Working memory is small and it decays fast. If retrieval is slow, you are still assembling word two when word six arrives, and at some point the buffer overruns and you lose the sentence — not one word of it, all of it. That is the experience of catching the first word and nothing else. It is not that the rest was too hard. It is that you never got to it.
This is also why the failure is so abrupt. Comprehension against speed is not a gentle slope; it holds up and then falls off a cliff, because the process is a queue and queues do not degrade gracefully.
Why more vocabulary does not fix it
The instinctive response is to learn more words. It is the wrong lever, and you can see why from the list above: your problem is at step 2, but it is not coverage, it is latency. You know the word. You cannot get to it fast enough.
A word you have met a hundred times in context is retrieved in a different way from a word you have drilled twenty times on a card. The flashcard gets you recognition; it does not get you automaticity. And listening needs automaticity, because there is no pause button in a conversation.
So adding a thousand words to your deck adds a thousand slow words. It raises your reading ceiling, where you control the pace. It does very little for the sentence at 3.6 words a second.
What does fix it
Three things, in rough order of how much they matter.
Volume at the speed you are failing at. You get faster at the rate you practise at. Hundreds of hours of slow, clear audio makes you very good at slow, clear audio. At some point the practice has to be at natural pace, which means real material, spoken by people who are not performing for learners — at a level you can nearly follow, which is the whole subject of comprehensible input.
Repetition of the same short piece. This is the part people skip, and it is the highest-value habit in listening practice. Take a clip of twenty to forty seconds, listen four times, read the transcript, listen twice more. The fourth listen is a completely different experience from the first — you are hearing the boundaries. Doing that for ten clips beats getting through an hour of new audio, by a wide margin, and it is one reason short-form video is a better training set than a podcast episode.
Reading along with the audio, then without it. A transcript turns an incomprehensible stream into something segmented, and hearing the stream while your eye follows the segmentation is how the ear learns where the joins are. The critical part is taking the transcript away afterwards. Reading along forever means you are reading, not listening.
Gradual is doing the work, not being gentle
None of the above says to jump straight to the fastest material you can find. Speech you catch nothing of is not training. The target is the same as anywhere else in learning: the edge of what you can do. Material you follow most of, at a rate slightly above the one you are comfortable with.
That edge moves, which is the annoying part — anything at your level now is beneath it in two months. It is why a fixed playlist of "beginner listening" stops working and why levels that are a one-off tag on a piece of content go stale. The rate at which a clip is spoken is a measured property of that clip; the rate you can follow is a moving property of you; and useful material is the overlap between the two.
We built the overlap into the product — every clip's speaking rate is measured from its transcript before anyone is shown it, and the feed uses it — but the principle stands on its own, and you can apply it by ear. When you find something you can follow about three-quarters of, stop looking for something better. That is the thing that is going to move you.
What to expect
The change, when it comes, is not gradual in the way you would like. People typically report months of no perceptible progress followed by a fairly sudden shift, where material that was a wall becomes merely hard. That is consistent with the queue model: nothing visible happens while retrieval times are falling, and then they cross the threshold where assembly can keep up and the whole thing works.
Which means the only real mistake is concluding, at month four, that it is not working. That is what it looks like while it is working.
Two things worth reading next, because they are the two levers that actually move this: how few words carry most of a language, which is where retrieval speed comes from, and which subtitles to use and when, which is where most listening practice quietly turns into reading practice.