Back to the site
Home/Blog/A1 to C2 explained by what you can actually watch

A1 to C2 explained by what you can actually watch

8 September 2026 · 6 min read

What the CEFR levels mean in practice, how to tell which one you are at without a test, and why a level tag on a video is usually somebody guessing.

On this page

  1. The six, in one line each
  2. Where people actually sit, and for how long
  3. Which level am I?
  4. The problem with a level on a piece of content
  5. What a model is still needed for
  6. Using levels without letting them use you

The Common European Framework of Reference gives six levels, A1 through C2. They are everywhere — on courses, on exams, on job listings, on the label of every "B1 podcast" in your feed — and almost nobody can say what they mean without looking them up.

The official descriptors are written for institutions. This is the version written for somebody deciding what to watch tonight.

The six, in one line each

A1 — you can catch prepared, isolated phrases. Greetings, numbers, prices, "where is the station". You understand when somebody is speaking slowly and directly to you and has agreed to be patient. You do not follow two other people talking.

A2 — you get the gist of simple, concrete speech about familiar things. Somebody describing their routine, ordering food, explaining where they live. You catch the topic and the main claim. The details escape.

B1 — you follow clear standard speech on familiar subjects. This is the big one, and the jump from A2 is the largest in the ladder. At B1 you can watch a video about something you already know about and come away with what it said. You cannot yet handle someone changing the subject unexpectedly, or being funny.

B2 — you follow most of what is said, including some argument and some abstraction. News, an interview, a vlog on almost any everyday subject. You still lose ground with heavy accents, fast overlapping speech and jokes that turn on wordplay.

C1 — you follow extended speech even when it is not clearly structured. Implicit meaning, tone, irony. You notice when somebody is being sarcastic. This is roughly "watch what you like, mostly without effort".

C2 — everything, including the parts native speakers find hard. Rare in practice among learners, and not a useful target for most people.

Where people actually sit, and for how long

Two honest observations that the ladder hides, because the rungs look evenly spaced and they are not.

A1 to A2 is fast. A few months of consistent work, in a language close to your own. The commonest few hundred words carry an enormous share of simple speech, so early progress feels rapid. See how many words you need for why.

B1 to B2 is where everybody stalls. It is a long plateau, it is where most learners have been for years, and the reason is that the remaining gains come from the flat part of the vocabulary curve plus a listening-speed problem that vocabulary does not fix. Most people who say "I'm intermediate" have been intermediate for a long time.

This is also why catalogue design matters more than it sounds. If you build a library of content and let it fill up with whatever native short-form video mostly is, you get a library that is almost entirely C1 — fast, idiomatic, full of references — and unusable by the people who have just started. Our own content targets are weighted deliberately towards the bottom of the ladder, because that is where learners are: roughly a third of each language's catalogue at A1, a third at A2, a quarter in the B range and the rest above it. Left to itself, the sweep produces the opposite distribution.

Which level am I?

You do not need a test. Take a piece of real material — a short video made for native speakers, not for learners — and watch it once, without subtitles.

  • You caught a handful of isolated words. A1.
  • You could say what it was about, but not what was said about it. A2.
  • You could summarise it in two sentences and you missed some details. B1.
  • You got nearly all of it, and lost only the jokes and a couple of fast bits. B2.
  • You did not notice you were doing it in another language until afterwards. C1.

Do it with three different clips on three different subjects, because your level is not one number — it is much higher on subjects you know about and lower on ones you do not, which no certificate records.

The problem with a level on a piece of content

Here is where the framework gets abused.

A level on a person is a summary of a broad ability, assessed over several skills. A level on a video is usually one of two things: a guess by whoever uploaded it, or a default that nobody ever changed. Both are close to useless, and the second is worse than nothing because it looks like information.

If a level tag on content is going to mean anything, it has to be derived from the content. Three properties of a clip can be measured from its transcript, without a model and without an opinion:

How fast it is spoken — words per second of actual speech, silence excluded, against a reference rate for that language. This is the wall beginners hit first and it varies enormously between clips that look identical on a thumbnail.

How far into the vocabulary it reaches — what share of the language's running vocabulary you would need to know to have met every word in the clip. Measured against a frequency corpus rather than a word list.

How fast new words arrive — words used for the first time per second of speech. A clip can use ordinary vocabulary and still be punishing if it introduces something unfamiliar every two seconds.

Take the harsher of the first two as the floor — a clip is as hard as its worst property, not its average — and adjust for the third. That gives a number per clip that is a measurement rather than a label, and it can be mapped back onto the CEFR bands so the tag on the clip means something.

The part worth emphasising: nothing about that asks a model how hard a clip is. Speaking rate and vocabulary reach are arithmetic over a transcript. Asking a language model to rate difficulty produces a plausible number with nothing behind it, and it will cheerfully rate the same clip differently twice.

What a model is still needed for

There is one thing about a clip you cannot count your way to, and it is what the clip is about. Topic is not in the numbers. Neither is whether the speaker is holding up the thing they are naming, whether the accent is regional, or whether the humour depends on something you would have to have grown up with.

So: level is derived and never asked, topic is asked and never derived, and a human reviewer can overrule either. That split is not a technicality — it is the difference between a library you can trust the labels on and a library where every tag is a guess with a confident face.

Using levels without letting them use you

Three practical rules.

Do not use your level as a filter. Work a little above it, always. The material that moves you is the material you follow about three-quarters of, and that is by definition above the rung you are comfortable at.

Do not wait to be "ready" for native content. You are ready now; you just need the easy end of it. Native content is not one difficulty. A person explaining a recipe slowly to a camera and two friends arguing about football are both native content and they are three rungs apart.

Ignore the certificate unless somebody is asking for it. For a job or a visa, take the exam. For learning, the level is a rough description of where you are, and the only number that changes anything is how many hours you spent understanding things this week.

Try it on a real clip

Vako is a vertical feed of real short-form video in Japanese, Spanish or Korean, with the transcript under the clip, the translation one tap away and every word one tap from its meaning. It is free to start, and it needs no account.

See how it works

Read next

  • Comprehensible input, and the part nobody tells you about getting some What comprehensible input actually means, why most beginners never find any, and a practical way to get it from video you would have watched anyway.
  • How long it really takes to learn a language, at fifteen minutes a day The honest arithmetic behind every study plan: what the FSI hour estimates mean, why they are not your number, and what a daily habit actually buys.
  • Why you can read the sentence and still not catch a word of it Native speakers are not mumbling. They are speaking at two to three times the rate of your textbook audio, and the fix is not more vocabulary.

© 2026 Vako

Home · Blog · Terms · Privacy · Delete your account · hello@abhishekbr.com

Videos play through YouTube’s official player and remain on YouTube. Vako is not affiliated with, endorsed by or sponsored by YouTube, Google, Apple or any creator whose work appears in the app.