Every few months someone rediscovers that people learn languages by understanding things, posts about it, and sets off a week of arguments. The idea itself is nearly fifty years old and it is not really in dispute. What is in dispute — and what almost nobody writes about — is the boring logistical problem underneath it: where an adult beginner is supposed to find several hundred hours of material they can nearly understand, in a language they cannot yet understand.
That gap is the whole subject of this post.
What the term actually means
Comprehensible input is language you meet — heard or read — that you understand most of, but not all. The phrase comes from Stephen Krashen, who argued in the early 1980s that people acquire a language by understanding messages slightly beyond their current level, and shorthanded that as i+1: what you have, plus one step.
Two halves, and people drop one of them constantly.
Drop the comprehensible and you get the beginner who puts on a film in Portuguese with no subtitles and lets it run for two hours. That is not input; it is noise with occasional words in it. Nothing is being understood, so nothing is being acquired, and the only thing the two hours reliably produces is the belief that you are bad at languages.
Drop the input and you get the opposite failure, which is far more common because it feels like work: six months of grammar tables, conjugation drills and gap-fill exercises, at the end of which you can name the past subjunctive and still cannot follow two people ordering coffee. You have learned about the language rather than having met any of it.
The claim is not that study is useless. It is that study is the small part, and the large part is hours — a lot of hours — of understanding real messages.
Why the +1 is harder to hit than it sounds
Here is the problem with i+1 as advice: it is a description of a state, not an instruction. It tells you what good material feels like once you have found it. It does not tell you how to find it, and for most languages the honest answer is that almost nothing you can reach is at your level.
Consider what is actually available to somebody four weeks into Japanese.
Textbook audio is comprehensible and is not really input — it is scripted, spoken at an unnatural pace by voice actors who have been told to enunciate, and it contains the grammar point of the chapter roughly nine times. Podcasts for learners are better and there are some excellent ones, but a beginner exhausts the beginner tier in weeks and the next tier up is a cliff. Native content — television, film, YouTube, anything a native speaker would choose — is not at i+1. It is at i+40.
So the practical situation is: a narrow band of material made for learners, which runs out, and then an enormous ocean of material that is far too hard, with a gap in between where most people quietly stop.
The graded-reader world solved this for reading decades ago, by writing books at controlled levels. Nobody has solved it for listening at scale, because you cannot write a graded YouTube channel — the content has to already exist, and then somebody has to know how hard each piece of it is.
The thing that makes short video unusually well suited
Short-form video turns out to sit in an odd and useful place.
It is short. A forty-second clip is a unit you can watch four times without resenting it. Repetition is where a lot of the acquisition happens, and nobody rewatches a forty-minute episode four times.
It is real, but it is performed for clarity. A creator explaining how they make their breakfast is speaking naturally, but they are also speaking to a camera, to strangers, with a point to make in thirty seconds. That is measurably slower and more clearly articulated than two friends talking in a bar, and it is still not a textbook.
The visual context carries a lot of the meaning. When somebody holds up an onion and says the word for onion, the word is comprehensible in the strict sense: you understood the message. Video does for free what a textbook does with a glossary.
There is an effectively unlimited supply of it, in every language with a young population, on every subject a person could be interested in.
The catch is the same as always. Somewhere in that unlimited supply there are clips a particular beginner could follow, and there is no way to find them by hand. Watching thirty clips to find the one that was at your level is thirty clips of failure to buy one clip of learning.
Making the haystack sortable
This is where the problem stops being pedagogical and starts being ordinary engineering: if you have the transcript of a clip, you can measure it, and if you can measure it you can sort it.
Three things are worth measuring, and all three are arithmetic rather than opinion.
How fast it is spoken. Not the clip's length — the rate of actual speech, with the silence between lines taken out. This is the wall beginners hit first, and it varies enormously between clips in a way nothing on the thumbnail tells you.
How far into the language's vocabulary it reaches. Rank every word in the language by how often it is actually said, then ask what share of the words in this clip sit inside the commonest few hundred. A clip built from very common words is followable long before one that is not, whatever its subject.
How fast new words arrive. A clip can use ordinary vocabulary and still be punishing if it introduces an unfamiliar word every two seconds. Density is its own axis.
Put those three together and you have a number per clip that means something, and the ability to hand somebody the material at their level rather than telling them to go and find some. That is what we build: the measurement runs over every clip's transcript before anyone is shown it, and the feed is ordered by it. You can read more about how a clip is graded and about how few words carry most of a language.
What to do with this, starting tonight
You do not need any of our machinery to apply the idea. You need three habits.
- Pick material by how much of it you understand, not by how much you like the subject. Interest matters, but it is the tiebreaker, not the filter. If you are catching under half, it is too hard, and no amount of determination changes that.
- Rewatch. The second pass through a clip you half-understood is worth more than a first pass through a new one. On the second pass you are hearing the shape of it rather than scrambling.
- Look things up sparingly, and only after. Stopping on every unknown word turns input into study and kills the flow that the whole method depends on. Watch it through, then check the two or three words that stopped you, then watch it again.
And be honest about volume. The people who make this method work are not doing twenty minutes on a Tuesday. They are doing an hour a day for a year or more, and the reason they can sustain it is that they are watching things they would have watched anyway. That is the real argument for short video over a textbook: not that it is better pedagogy in the abstract, but that it is the only form of input most adults will actually consume in the quantities the method requires.
The honest limits
Two things the strong version of the input hypothesis gets wrong, or at least overstates.
Some explicit study pays for itself. Twenty minutes learning how your target language marks its past tense will make hundreds of hours of listening more comprehensible. The evidence for a pure no-study approach is much weaker than its advocates suggest; the evidence that study alone is insufficient is overwhelming.
Output is not automatic. Understanding a great deal will make you a good listener and a fast reader. Speaking is a separate motor skill, and it comes from speaking. You can postpone it — many people do, deliberately, for a year — but you cannot skip it.
What the input case gets right is the proportion. If your last six months of study contained more hours of exercises than hours of understanding real language, the balance is wrong, and fixing that single ratio will do more for you than any other change you could make.