Back to the site
Home/Blog/Subtitles when learning a language: which ones, and when to turn them off

Subtitles when learning a language: which ones, and when to turn them off

28 August 2026 · 5 min read

English subtitles, target-language subtitles, or none. What each one trains, what the evidence suggests, and a rule that survives real use.

On this page

  1. What each one actually does
  2. The rule that works
  3. Why a transcript beats a subtitle track
  4. The specific mistake to avoid
  5. What to do when nothing is comprehensible
  6. The short version

Put on a video in the language you are learning and you immediately face a decision that people have strong opinions about and very little basis for: English subtitles, subtitles in the target language, or nothing at all.

The honest answer is that all three are useful and they train different things, and the common advice — "turn them off, they are a crutch" — is wrong in a way that costs beginners a lot of time.

What each one actually does

English subtitles. Your eyes go to the English. This is not a matter of willpower; reading in your first language is close to automatic, and an automatic process wins against an effortful one every time. The eye-tracking work on subtitled video consistently shows that viewers read the subtitles when they are there, and that when those subtitles are in a language you read fluently, very little attention is left for the audio.

What English subtitles are good for: enjoying something, following a plot so you are not lost, and checking your understanding after you have tried without them. What they are not good for: any pass where you intend to be learning from the audio.

Target-language subtitles. These are the interesting case, and the evidence is generally favourable. They do something specific and valuable: they show you where the word boundaries are. Most of what makes a second language sound like an undifferentiated stream is a segmentation problem, and subtitles in the target language solve it visually while your ear is listening to the unsegmented version. That pairing is how the ear learns where the joins are.

They come with a real risk, which is that you stop listening and start reading. Reading is easier and you will drift into it if the pass goes on long enough.

No subtitles. The only one that trains the actual skill, and useless on material far above your level. Nothing is learned from an hour of noise.

The rule that works

Use all three, in a fixed order, on the same short piece.

  1. Nothing. One pass, cold. This is your measurement, and it is also the pass that forces your ear to do the work.
  2. Nothing, again. You will get more the second time purely from knowing what is coming. This is free and almost everybody skips it.
  3. Target-language text, no audio. Read it. Find out what was actually said. Most of what you missed will turn out to be words you know, run together — which is a completely different diagnosis from "I do not know enough words", and it points at different work.
  4. English, briefly, only for the bits that still do not make sense. Not the whole thing. The two or three places you are still stuck.
  5. Target-language text with audio, twice. This is the pass that does the work: the sound and the segmentation together.
  6. Nothing, once more. The rep that counts.

Six passes over thirty seconds of video. It takes about ten minutes and it is worth more than an hour of new material with English on the screen.

Why a transcript beats a subtitle track

There is a practical difference between a subtitle track and a transcript, and it is larger than it sounds.

A subtitle track is timed to the video and disappears. You cannot dwell on a line, you cannot look at the previous one, and you certainly cannot tap a word. If the line goes past and you did not get it, your options are to rewind or to lose it.

A transcript under the video is a text you control. You can read ahead, read back, stop on a word, and — if the transcript knows what its words are rather than being a wall of characters — get the meaning of one of them without leaving the clip.

For a language that does not put spaces between words, this stops being a convenience and becomes the whole thing. A Japanese subtitle line is a continuous string; a transcript that has been split into words, with the reading above each one, is a different artefact doing a different job.

That is the specific thing we build: the clip plays, the transcript sits under it, the English translation is one tap away rather than on screen by default, and every word is tappable for its meaning in that sentence. The ordering is deliberate — the translation being a tap away rather than visible is what stops step 4 above from swallowing steps 1 through 3.

The specific mistake to avoid

The failure mode that burns the most hours is this one: watching a lot of content with English subtitles on, for months, and believing it is immersion.

It is not. It is watching foreign-language television. It is a pleasant thing to do and it will teach you a small amount incidentally, but the ratio of hours to learning is terrible, and the reason it persists is that it feels productive — you are understanding the story, so you are understanding, so surely you are learning.

The test is simple. Turn the subtitles off for thirty seconds. If comprehension drops to near zero, you were reading.

What to do when nothing is comprehensible

If a clip is so far above you that even the transcript does not help, the answer is not to add English subtitles and push through. It is to find easier material.

This is the hard part of the advice, because easier native material is not easy to find by hand — you cannot tell a clip's difficulty from its thumbnail, and thirty clips of failure to find one good one is thirty clips of discouragement. Difficulty is measurable from a transcript, though: speaking rate, how far into the vocabulary it reaches, and how fast new words arrive. Something should be doing that sorting for you, whether it is a tool or a patient friend with the same target language.

The argument for why that matters more than any subtitle setting is in comprehensible input and how to get some, and the reason "just turn them off and push through" fails is in why you can read the sentence and still not catch a word.

The short version

  • English subtitles: for enjoyment, and for checking after the fact. Never for a listening pass.
  • Target-language text: for segmentation, and it is genuinely valuable. Pair it with audio, not instead of audio.
  • Nothing: the only pass that trains the skill, and it has to be on material you can nearly follow.

And the rule underneath all three: the number of passes matters more than the setting. Anybody arguing about subtitles while watching each clip exactly once is optimising the wrong variable.

Try it on a real clip

Vako is a vertical feed of real short-form video in Japanese, Spanish or Korean, with the transcript under the clip, the translation one tap away and every word one tap from its meaning. It is free to start, and it needs no account.

See how it works

Read next

  • Why you can read the sentence and still not catch a word of it Native speakers are not mumbling. They are speaking at two to three times the rate of your textbook audio, and the fix is not more vocabulary.
  • Comprehensible input, and the part nobody tells you about getting some What comprehensible input actually means, why most beginners never find any, and a practical way to get it from video you would have watched anyway.
  • How to learn Japanese from YouTube Shorts without wasting your time A practical method for using short-form Japanese video as real study: what to watch, what order to do it in, and the three traps that make it feel useless.

© 2026 Vako

Home · Blog · Terms · Privacy · Delete your account · hello@abhishekbr.com

Videos play through YouTube’s official player and remain on YouTube. Vako is not affiliated with, endorsed by or sponsored by YouTube, Google, Apple or any creator whose work appears in the app.