Verbamor Download the appGet the app
All posts Download on the App Store
Memory science

The production effect: why saying it out loud beats reading it

Reading a word aloud made it far easier to recognise later. Then the same lab ran the version where every word was read aloud, and the whole advantage vanished.

You're on the sofa with 30 cards. You tap through them with your mouth shut. Every single one feels known.

Acaso comes up. Maybe, perhaps. You swipe on.

Then Thursday comes and your tutor asks you something. The word is gone.

The difference between thinking a word and saying it

Say a word out loud while you study it and you'll remember it better than one you only read. That's the production effect. It has a name and a pile of experiments behind it.

MacLeod and colleagues put it through 8 recognition experiments in 2010. In Experiment 1B, out of every 100 studied words, people correctly recognized about 20 more of the ones they'd spoken. That's 20 percentage points for the cost of moving your lips.

Older studies they review landed in roughly the same territory, somewhere in the teens to mid-20s. So the size holds up across labs.

Every one of those gains is a comparison inside one person's list. Reading the whole deck aloud pays, and reading half of it aloud pays more, because the spoken half has something to stand out against.

The words you speak get their edge from the words you don't.

So the ruling for your 30 cards. Say roughly 15 out loud and leave the other 15 silent, and the 15 you speak stand out against the 15 you don't. Speak all 30 and you flatten the list back to the smaller pure-list effect.

Say each of those 15 from a blank prompt. Cover the answer, say the Spanish before you flip the card, then check. Pick the 15 from the words you keep dropping in conversation.

And say it properly. Full voice, the volume you'd use ordering a coffee, with air moving through your vocal cords and your ears picking up the sound. The memory gets its edge from a real motor act and a real sound, so the more of that you cut, the less edge you get.

Thursday is the first time your mouth will make that sound. Make it Tuesday instead.

What MacLeod measured

Saying a word aloud beat reading it silently by somewhere between .064 and .204, depending on the experiment. Students studied a list (MacLeod and colleagues, 2010), some words aloud and some silently, then took a recognition test.

The top of that range comes from Experiment 1B. The paper puts it flatly: "Hits for words read aloud exceeded by .204 those for words read silently during study."

A hit rate is the share of studied words you correctly pick out later. Study 20, and 4 more land in the "yes, I saw that" pile.

Experiment 1A came in at .124, and from there it slides.

  • Experiment 3A: .133
  • Experiment 7: .096. Here the silent items were generated, so students had to work the word out.
  • Experiment 8: .064. Here the comparison task was a semantic judgment.

The harder the silent condition works, the less speaking buys you. Against a semantic judgment it wins 1 word in 16. Pulling the word out of memory cold does more, and that's the testing effect. Speaking is a garnish on retrieval practice.

Volume decides which number you get. MacLeod's team went after this directly, and the ranking is blunt: full voice wins, while mouthing and whispering land closer to silence than to speech.

So say the words at the volume you'd use to answer a question in a kitchen. Once each.

Older work reported bigger numbers, and MacLeod's paper reviews both: Conway and Gathercole found a 15 to 25 percent aloud-over-silent advantage, Gathercole and Conway 14 to 20 percent. I trust the smaller numbers because the paper hands me a separate figure per experiment. The older percentages arrive as a range with the comparison conditions folded up inside them. When a number gets more careful and gets smaller, I plan around the smaller one.

So I'd plan around .05, the recognition figure of roughly .1 cut in half because I'm guessing at a job nobody measured. On 6 spoken words that's a fraction of a word per session.

Read everything aloud and it stops working

MacLeod ran Experiment 2 with 2 separate groups. One read every word aloud, the other every word silently.

The production effect disappeared. "The main effect of study condition was not significant, F(1, 28) = 1.97, MSE = 0.018, p > .15," the authors write, "nor was the interaction (F < 1), together indicating no production effect in the between-subjects situation." On hit rates alone: t(28) = 0.95, p > .30.

The aloud group did recognize .062 more of the studied words. They also falsely claimed .070 more they'd never seen. The gain and the cost cancel.

The authors say the effect needs a mixed list, studied by one person, calling it "restricted to within-subject, mixed-list designs."

That's 30 people in one experiment, so I wouldn't call it settled. Fawcett (2013) pooled the between-subject studies and found a modest effect in pure lists, g = 0.37, on recognition. Revisiting it in 2023, his team saw recognition survive and recall vanish.

This is the logic behind desirable difficulty. The work has to be uneven to leave a mark.

All 8 experiments are recognition tests: a word flashes on screen and you answer yes or no, did I study this?

Your Thursday job is harder. Your tutor says "napkin" and you produce la servilleta from nothing, with somebody waiting. Anyone quoting .204 for vocabulary recall is borrowing a number from a different task.

The recall evidence is thinner. Conway and Gathercole switched their third experiment to free recall and the aloud advantage held at a similar size. One experiment, though.

I'd still run the ratio. Saying a word aloud costs you about a second, so the downside is tiny even if the effect turns out to be recognition-only.

Speaking your vocabulary buys you memory for the words themselves. Building sentences and understanding what someone says back need retrieval practice and real listening.

Those quiet words pay for the loud ones. The silent group in Experiment 2, who never spoke all session, hit .062 higher than the mixed-list silent words did. You lift the loud words a lot and sand a little off the quiet ones.

That trade only works if you pick the loud ones honestly.

Why the mouth adds something the eye does not

The spoken word arrives at test time carrying extra baggage. You heard yourself, and your jaw moved.

When you're asked whether you studied this one, the baggage answers for you. It only answers because the words around it came empty-handed.

That feeling of having handled the word is easy to misfile. If every word feels handled, it stops sorting them.

So say 5 to 10 words per lesson out loud, and speak no more than a quarter of what's in front of you. That's the ratio MacLeod's within-subject experiments ran on: half of a 20-item list spoken, half silent.

Pick the ones your tutor corrected you on. A correction is already the item you got wrong under pressure, which beats anything you'd choose by feel at 9pm.

Some weeks the corrections don't reach 5. Fill the gap with the phrases you reached for and couldn't finish, and stop at 10 all the same. Some weeks you'll count 22, so keep the ones corrected more than once and take the first 10.

Read the words you leave silent properly. Skim them and you flatten the contrast just as thoroughly as speaking all 20 does.

MacLeod ran 20-item lists with undergraduates. Nobody has run the version where a learner speaks 8 of 40 phrases and gets tested six weeks later.

Roberts and colleagues (2024) had people read passages aloud or silently. Production lifted recall of what they'd read while comprehension sat flat.

Production puts a bookmark in an item. A bookmark marks a location, and the meaning of the page is still your job.

When your tutor writes out why the subjunctive shows up after para que, read that silently and ask your question about it out loud. A rule sinks in when you interrogate it.

Desirable difficulty and the testing effect both run on hard retrieval. Production runs on items sticking out from their neighbors. The shapes rhyme and the machinery differs, so don't expect spoken phrases to do what a retrieval drill does.

How to use this on Thursday

My deck has 30 cards in it. 10 of them I keep missing.

  • Sort the deck before you open your mouth. Run all 30 silently. Any card you get wrong, or sit on for more than 3 seconds, joins the loud pile. Re-sort every session, because a card you've finally learned has to go quiet again.
  • I picked the numbers here myself. MacLeod's paper gives you the mixed-list rule and stops there. I picked 3 seconds because past that you're reconstructing the word. Move them if your deck says otherwise.
  • A loud card goes quiet after 2 clean sessions. Clean means the right answer inside 3 seconds without stalling. 1 good session is luck. 2 in a row is the card.
  • A quiet card leaves the 30 after 4 clean sessions. Drop a new word in its place, and the deck stays at 30 with roughly 10 loud. If the loud pile climbs past 12, stop adding new words.
  • Give your voice to the 10 you keep missing. The other 20 stay quiet. They're the background the loud ones stand out against.
  • Produce before you flip. Cover the answer, say the Spanish out loud, then reveal the back. See the testing effect for why the guess itself does the work.
  • Whisper when a full voice is impossible. MacLeod's comparison was aloud against silent, so a whisper sits outside what they measured. I think it keeps most of the benefit. Treat that as my extrapolation.
  • Don't read the whole deck aloud. You'd spend the extra minutes and land where you started, running MacLeod et al. (2010) Experiment 2 on yourself.
  • Speak the vocabulary, write the grammar. Roberts et al. (2024, Memory & Cognition) found the boost lives in memory for the words themselves, while reading comprehension sits flat. Give Thursday's subjunctive rule a page instead, and write 4 sentences with it.

Speaking buys you something on top of the memory. Your mouth has to learn the shape of a French r, and you build that by repeating it until it stops feeling strange.

Verbamor builds in most of this. Every card is a sentence with the word cut out, so you produce the answer before you see it, and it carries a native recording of the full sentence to check against.

The one part I can't automate is you opening your mouth.


Sources

Saying a word aloud beats reading it silently, by .064 to .204 depending on the experiment

Speak the whole list and the effect dies, because the gain in hits gets eaten by false alarms

  • MacLeod, C. M., Gopie, N., Hourihan, K. L., Neary, K. R., & Ozubko, J. D. (2010). The production effect: Delineation of a phenomenon. Journal of Experimental Psychology: Learning, Memory, and Cognition, 36(3), 671-685. Experiment 2, the between-subjects null.

Older studies put the aloud advantage higher, in the teens to mid-20s as a percentage

  • Conway & Gathercole, and Gathercole & Conway, both reviewed inside MacLeod, C. M., Gopie, N., Hourihan, K. L., Neary, K. R., & Ozubko, J. D. (2010). The production effect: Delineation of a phenomenon. Journal of Experimental Psychology: Learning, Memory, and Cognition, 36(3), 671-685.

Every claim above is a recognition result, so the recall numbers are my extrapolation

  • MacLeod, C. M., Gopie, N., Hourihan, K. L., Neary, K. R., & Ozubko, J. D. (2010). The production effect: Delineation of a phenomenon. Journal of Experimental Psychology: Learning, Memory, and Cognition, 36(3), 671-685. All 8 experiments test recognition.

Every load-bearing claim Verbamor makes is traced to its paper on the research page.

Cards you answer out loud.

Verbamor asks you to produce the word into a blank, then plays a native speaker saying the whole sentence.

Download on the App Store