Verbamor Download the appGet the app
All posts Download on the App Store
Tactics

Why you understand the show but not the person at the counter

You just followed 50 minutes of television. Then a barista says one sentence and you're gone. The counter isn't using harder words. It's running about 50 words a minute faster with the edges rubbed off.

Last month I watched a full episode of a Spanish show without subtitles and understood most of it. Felt great. The next morning a woman at a bakery in Madrid asked me something about six words long and I had nothing. Not a partial guess. Nothing.

I made her repeat it twice. On the third pass it broke apart into words I'd known for a year. She'd asked whether I wanted the bread sliced.

That gap has almost nothing to do with your vocabulary. The counter is running about 50 words a minute faster than the television, and at that speed the words stop having edges. Both of those have been measured, and one of them moves in about a month.

The counter is just faster

Start with raw speed, because it's the part people underrate.

Steve Tauroza and Desmond Allison ran the classic measurement in 1990. They timed British English across 4 kinds of speech: radio monologue, interviews, lectures, and casual conversation. Everybody quotes their average of 170 words a minute. The interesting part is how far apart the categories sat.

Wang Li redid the study in 2020 for the Academic Journal of Humanities and Social Sciences, running the same method on 200 British speakers and printing the table category by category. Lectures came in at 173 words a minute. Casual conversation came in at 222.

Wang (2021) · British English, words per minute
0
a lecture, the slowest category measured
0
casual conversation, the fastest
0
the top of the conversation range
Wang analyzed 120 minutes of speech from 200 native speakers, replicating Tauroza and Allison's 1990 method. Conversation was the fastest of the 4 categories.

That's 49 extra words a minute, or about one extra word per second.

You don't have a spare second. Your comprehension is running near its limit already, and conversation hands you 28% more words in the same window.

You didn't get slower. The audio got faster, and you were already at your ceiling.

Wang measured British English, so take 222 as the shape of the thing rather than your language's number. The ranking is what I'd expect to travel. A lecture is one prepared person trying to be clear, and a conversation is 2 people racing each other to the end of the sentence.

Nobody puts spaces between words

Speed alone would be survivable. The second problem is what speed does to the edges of words.

Written language has spaces. Speech has none. The sound coming at you is one continuous stream, and your brain has to decide where each word stops. Researchers call that job lexical segmentation, which is just cutting the stream into words.

In your first language you never notice this, because you've been doing it since before you could read. In a new language you're doing it consciously, on audio that's moving too fast to think about.

And the words don't help. Fast speech mashes them together. Endings drop, vowels flatten, and two words fuse into one shape that isn't in your head. English turns "what do you" into whaddya, and every language has its own version of that.

Here's how badly it goes. Sheppard and Butler tested 77 learners with a task that stops the audio mid-conversation and asks you to write down the last few words you heard. Nothing to read, no trick. The learners got 67% of the words right.

Two thirds, with the sound still ringing in their ears. Hand those same learners the same sentences on paper and they'd read them without trouble, because the vocabulary was never the hard part.

One learner in that study heard "don't always notice" and wrote down "don't always know this." Every sound is roughly there. The seams are in the wrong places.

Kriss Lange and Joshua Matthews ran the same kind of test on 130 Japanese students in 2020. They stacked the deck toward easy on purpose. 97% of the words they asked for were from the most common 1,000 in English.

Even then, how well a student cut the stream predicted their score on a real listening exam. Segmentation and vocabulary together accounted for about a third of the gap between students.

The gap between knowing and hearing
On the page attracts investment Both words are ordinary. Read them and you understand them instantly.
In the ear a tax investment What learners actually wrote down. The boundary moved, and the sentence stopped making sense.
A real transcription error recorded by Matthews and O'Toole (2015), cited in Lange and Matthews (2020).

Look at what happened there. Nobody failed to know "attracts". They heard the sounds correctly and cut them in the wrong place, and out came a phrase that means nothing.

Then they lost the next sentence too, because they were still trying to make "a tax investment" work. One bad cut takes the rest of the paragraph with it. That's the counter, exactly.

You're losing the small words

There's a pattern in which words go missing, and it's the opposite of what you'd guess.

John Field tested this in a 2008 paper in TESOL Quarterly called Bricks or Mortar. He played sentences to intermediate learners and looked at which words they recovered. Content words, the nouns and verbs, came back fine. Function words did much worse: the articles, prepositions, and pronouns.

Field found this held across learners with different first languages and different proficiency levels, which is what makes it interesting. It isn't a quirk of one group.

The reason is physical. Small words are short and unstressed, so they're the first thing a fast talker crushes. They're also where the grammar lives. Lose them and you're left holding a pile of nouns with no idea how they connect.

You know this feeling. You caught "bread", you caught "want", and you still couldn't answer, because the shape of the question was in the words you didn't catch.

Field also argues that most listening breakdowns are bottom-up, meaning they start with a misheard word rather than a failure to follow the topic. Which is worth knowing, because "listen for the gist" advice aims at the wrong layer entirely.

The show was carrying you

So why did the episode go fine?

The show was doing a lot of work for you, quietly. You could see the actor's face. You knew the plot, the characters, and roughly what this scene was for. When a line got past you, the next shot told you what it meant.

Scripted dialogue is also cleaner than real talk. It's written, then performed by people trained to be understood, then mixed by an engineer. Research comparing scripted and spontaneous speech finds the spontaneous stuff varies far more in speed, speeding up and slowing down inside a single sentence.

The bakery gave me none of that. No plot, one unknown speaker, no visual clue, and a real answer expected in about 2 seconds. Same language, most of the scaffolding removed.

Which means the hour of television was real practice, and it trained something narrower than you'd hope. If you want more out of that hour, the subtitle setting changes what your ear is doing, and picking a show at your level decides whether the plot is helping you or hiding the gap.

Training the ear you actually need

The encouraging part: this responds to practice fast, and the practice is boring in a useful way.

J. D. Brown and Ann Hilferty ran the study I keep coming back to, in 1986. They took 32 students and split them in two. One group got 10 minutes a day on reduced forms, the squashed versions of ordinary phrases. The other spent the same 10 minutes on minimal pairs.

After 4 weeks, the reduced-forms group went from understanding 35% of those sentences to 61%. The group doing minimal pairs didn't move the same way.

26 points in a month, on 10 minutes a day. It's a small study from 1986 and I'd like to see it run again with 300 people. But the direction matches everything above it, and the cost of testing it yourself is 10 minutes.

Here's what to actually do, and the first one is the whole game:

  • Transcribe 30 seconds a day, word for word. Take unscripted audio, a podcast where 2 people interrupt each other. Write down every word, including the small ones. Replay as much as you want, then check against a transcript. The gaps you find are your real listening level, and they'll be humbling.
  • Work the phrases that beat you. When a phrase defeats you, it's usually 3 known words fused together. Look at the transcript, find the 3 words, then play that second on loop until you can hear each one. Now say the fused version at full speed yourself, 10 times, because producing a form makes it easier to recognize later.
  • Ask your tutor to talk at full speed for 5 minutes. Most tutors slow down for you automatically, which is kind and leaves you unprepared for everyone else. Ask them to speak to you the way they'd speak to a friend, then work through what you missed.
  • Use audio without pictures. Podcasts and phone calls strip the visual scaffolding out, which is the point. If a show is the only thing you'll stick with, listen to an episode you've already seen with the screen off.

Do the transcription one even if you skip the rest. It's the only exercise that shows you the difference between the words you know and the words you can hear, and that difference is the entire problem.

Verbamor sits to the side of this problem. It handles the vocabulary half, taking the words from your lesson and scheduling them before you forget them, and the audio on every card gets you the word's sound in careful speech. You have to know a word before you can hope to catch it. But no flashcard will teach your ear to cut a fast stream into pieces. That part comes from the transcription drill and a lot of messy audio.

The bakery question, by the way, was ¿se lo corto? Three words, and I knew all three of them. At the speed she said it, I couldn't find the seams.


Sources

Conversation is the fastest speech

Cutting the stream into words

The small words go first

Reduced forms can be taught

Every load-bearing claim Verbamor makes is traced to its paper on the research page.

Know the word before you try to hear it.

Verbamor turns your lesson vocabulary into scheduled recall, with native audio on every card, so the word is already yours when someone says it fast.

Download on the App Store