Back to the site
Home/Blog/Why you can read the sentence and still not catch a word of it

Why you can read the sentence and still not catch a word of it

18 September 2026 · 7 min read

Native speakers are not mumbling. They are speaking at two to three times the rate of your textbook audio, and the fix is not more vocabulary.

On this page

  1. The number your textbook audio is hiding
  2. What speed actually does to comprehension
  3. Why more vocabulary does not fix it
  4. What does fix it
  5. Gradual is doing the work, not being gentle
  6. What to expect

There is a specific and very deflating moment that arrives about a year into a language. You can read a newspaper paragraph with a dictionary. You know several thousand words. You can construct a sentence with a relative clause in it and get the agreement right. Then somebody in a video says one short sentence at ordinary speed and you catch the first word and nothing else.

This is not a vocabulary problem and it is not a talent problem. It is a rate problem, and it is worth understanding precisely, because the wrong diagnosis leads to months of the wrong work.

The number your textbook audio is hiding

Speech has a measurable rate. For languages that write spaces between words, the natural unit is words per second; for Japanese or Chinese, which do not, you have to count characters instead, because "word" is not a thing the writing system marks.

The rates that matter are not the ones you have been listening to. Course audio is recorded by voice actors who have been asked to be intelligible, and it typically runs somewhere near the bottom of the natural range. Ordinary conversation between two native speakers who are not thinking about it runs much faster — in most languages close to double.

These are the reference rates we use internally to grade clips, in the unit each language is counted in:

LanguageDeliberate deliveryOrdinary conversation
Spanish2.2 words/sec3.6 words/sec
French2.1 words/sec3.4 words/sec
German1.9 words/sec3.0 words/sec
Korean1.3 words/sec2.3 words/sec
Japanese3.2 chars/sec5.6 chars/sec

Two things to take from that table.

The first is the ratio. In every one of those languages, ordinary conversation is roughly 1.6 to 1.8 times the rate of careful, deliberate speech. If your listening practice has all been at the deliberate end, you have been training on a signal that is most of a second per sentence slower than the one you are failing at.

The second is that the numbers are not comparable across languages. Spanish at 3.6 words a second and German at 3.0 are not telling you that Spanish speakers talk faster in any meaningful sense — German packs far more into each word. A single threshold across languages would mean "gentle" in one and "impossible" in another, which is exactly why anything that claims to grade content has to hold a reference per language rather than one number for everyone. Ours are informed estimates rather than laboratory measurements, and we say so; what matters for grading is the ratio between them, which is stable.

What speed actually does to comprehension

Here is why an extra word per second is so much worse than it sounds.

Understanding speech in a second language is not one process. It is at least three, running at once:

  1. Segmentation. Finding where one word stops and the next starts. In your first language this is automatic and you are not aware it is happening. In a second language it is work, and speech contains almost none of the gaps you imagine it does — the spaces you "hear" between words are mostly reconstructed by a brain that already knows the words.
  2. Retrieval. Getting from a sound to a meaning. For a word you know well this is instant. For a word you learned from a flashcard three weeks ago it might take a full second, and a full second is three or four words of speech.
  3. Assembly. Holding the pieces in working memory long enough to build the sentence.

The third one is where it collapses. Working memory is small and it decays fast. If retrieval is slow, you are still assembling word two when word six arrives, and at some point the buffer overruns and you lose the sentence — not one word of it, all of it. That is the experience of catching the first word and nothing else. It is not that the rest was too hard. It is that you never got to it.

This is also why the failure is so abrupt. Comprehension against speed is not a gentle slope; it holds up and then falls off a cliff, because the process is a queue and queues do not degrade gracefully.

Why more vocabulary does not fix it

The instinctive response is to learn more words. It is the wrong lever, and you can see why from the list above: your problem is at step 2, but it is not coverage, it is latency. You know the word. You cannot get to it fast enough.

A word you have met a hundred times in context is retrieved in a different way from a word you have drilled twenty times on a card. The flashcard gets you recognition; it does not get you automaticity. And listening needs automaticity, because there is no pause button in a conversation.

So adding a thousand words to your deck adds a thousand slow words. It raises your reading ceiling, where you control the pace. It does very little for the sentence at 3.6 words a second.

What does fix it

Three things, in rough order of how much they matter.

Volume at the speed you are failing at. You get faster at the rate you practise at. Hundreds of hours of slow, clear audio makes you very good at slow, clear audio. At some point the practice has to be at natural pace, which means real material, spoken by people who are not performing for learners — at a level you can nearly follow, which is the whole subject of comprehensible input.

Repetition of the same short piece. This is the part people skip, and it is the highest-value habit in listening practice. Take a clip of twenty to forty seconds, listen four times, read the transcript, listen twice more. The fourth listen is a completely different experience from the first — you are hearing the boundaries. Doing that for ten clips beats getting through an hour of new audio, by a wide margin, and it is one reason short-form video is a better training set than a podcast episode.

Reading along with the audio, then without it. A transcript turns an incomprehensible stream into something segmented, and hearing the stream while your eye follows the segmentation is how the ear learns where the joins are. The critical part is taking the transcript away afterwards. Reading along forever means you are reading, not listening.

Gradual is doing the work, not being gentle

None of the above says to jump straight to the fastest material you can find. Speech you catch nothing of is not training. The target is the same as anywhere else in learning: the edge of what you can do. Material you follow most of, at a rate slightly above the one you are comfortable with.

That edge moves, which is the annoying part — anything at your level now is beneath it in two months. It is why a fixed playlist of "beginner listening" stops working and why levels that are a one-off tag on a piece of content go stale. The rate at which a clip is spoken is a measured property of that clip; the rate you can follow is a moving property of you; and useful material is the overlap between the two.

We built the overlap into the product — every clip's speaking rate is measured from its transcript before anyone is shown it, and the feed uses it — but the principle stands on its own, and you can apply it by ear. When you find something you can follow about three-quarters of, stop looking for something better. That is the thing that is going to move you.

What to expect

The change, when it comes, is not gradual in the way you would like. People typically report months of no perceptible progress followed by a fairly sudden shift, where material that was a wall becomes merely hard. That is consistent with the queue model: nothing visible happens while retrieval times are falling, and then they cross the threshold where assembly can keep up and the whole thing works.

Which means the only real mistake is concluding, at month four, that it is not working. That is what it looks like while it is working.

Two things worth reading next, because they are the two levers that actually move this: how few words carry most of a language, which is where retrieval speed comes from, and which subtitles to use and when, which is where most listening practice quietly turns into reading practice.

Try it on a real clip

Vako is a vertical feed of real short-form video in Japanese, Spanish or Korean, with the transcript under the clip, the translation one tap away and every word one tap from its meaning. It is free to start, and it needs no account.

See how it works

Read next

  • Subtitles when learning a language: which ones, and when to turn them off English subtitles, target-language subtitles, or none. What each one trains, what the evidence suggests, and a rule that survives real use.
  • Comprehensible input, and the part nobody tells you about getting some What comprehensible input actually means, why most beginners never find any, and a practical way to get it from video you would have watched anyway.
  • How to learn Japanese from YouTube Shorts without wasting your time A practical method for using short-form Japanese video as real study: what to watch, what order to do it in, and the three traps that make it feel useless.

© 2026 Vako

Home · Blog · Terms · Privacy · Delete your account · hello@abhishekbr.com

Videos play through YouTube’s official player and remain on YouTube. Vako is not affiliated with, endorsed by or sponsored by YouTube, Google, Apple or any creator whose work appears in the app.