Verbamor Download the appGet the app
All posts Download on the App Store
Memory science

Interleaving has a negative effect size on words

The biggest meta-analysis of interleaving reports g = 0.42 overall. Split it by material and the words row reads g = -0.39. That's the row a vocabulary learner is standing in, and I published a post recommending the opposite.

Interleaving is one of the safest recommendations in learning science. Shuffle your practice instead of grouping it, and delayed test scores go up. Rohrer's math classrooms, Kornell and Bjork's paintings, usually filed alongside desirable difficulty. I've written that recommendation myself.

The meta-analysis everyone cites for it is Brunmair and Richter, 2019: 59 studies, 238 effect sizes, 158 samples. Overall Hedges' g = 0.42. That's the number in the abstract, and the one that gets quoted.

The same abstract has one more sentence in it. Sorted by learning material, the effect for words is g = -0.39. Negative. Blocking beats interleaving.

The row that flips

Here is Table 2 of that paper, reproduced in full, because I have not found another page that prints it.

Brunmair & Richter 2019, Table 2 · effect by learning material paintings · k 70 tastes · k 3 photographs · k 19 maths · k 45 artificial pics · k 63 texts · k 25 words · k 13 0.67 0.57 n.s. 0.35 0.34 0.31 0.21 n.s. -0.39 0
Hedges' g by material type. The words row sits on the other side of zero, where the bar points at blocking. n.s. marks the two rows whose confidence intervals cross zero.

The words row is not a rounding wobble near zero. Its 95% confidence interval runs [-0.64, -0.14], never touching zero, at p = .005. The overall interval is [0.34, 0.50]. Those two don't come close to overlapping.

Do the arithmetic the paper leaves on the table. Back the standard errors out of those intervals (0.041 overall, 0.128 for words) and the difference between them is 0.81, standard error 0.134. That's z = 6.05. The distance between "interleaving works" and "interleaving works on words" is itself one of the larger quantities in the paper, and it appears nowhere in the abstract.

One translation, since g is abstract. At g = 0.42 the average interleaved learner beats about 66% of blocked learners. At g = -0.39 they beat about 35%. Same technique, roughly mirrored around the middle.

Interleaving's own meta-analysis contains a row where interleaving loses, and it's the row with the words in it.

The honest caveat, which the authors state and I'll repeat: k = 13, about 5.5% of the paper, the thinnest row apart from tastes. The words estimate is also unusually consistent, at I² = 18.1% against 77.3% overall. Few, and in agreement. One wrinkle, since this post is about reading tables carefully: the paper prints that figure twice and the two disagree, 18.1% in Table 2 against 18.3% in the body text. Nothing turns on it, and I use the table because that is where the rest of the row comes from.

The obvious objection is that "words" stands in for something else: easy material, or unusual tasks, with the label taking the blame. Brunmair and Richter tested that. Their Table 3 regression controls for similarity within and between categories, complexity and familiarity at once, and the words penalty survives at b = -0.48 (SE 0.23, p = .014) against paintings. Familiarity mattered most of the material characteristics, b = 0.19 to 0.20, and words are the most familiar material this literature studies. The penalty outlives the control.

Two thirds of the evidence is pictures

Count the k column and the shape of the field falls out. Paintings 70, artificial pictures 63, naturalistic photographs 19: that's 152 of 238 effect sizes, 64% of everything, on visual materials. Every non-visual material put together comes to 86, and words are 13 of those.

So the general claim rests overwhelmingly on people learning to tell a Monet from a Seurat, and it gets applied to a flashcard deck because both activities involve categories. The materials where interleaving is best evidenced are the ones a language learner never touches.

One more number, on the general claim rather than the words row. A trim-and-fill correction for publication bias across the whole sample drops the overall estimate from 0.42 to g = 0.29, with 23 studies estimated missing: a 31% haircut on the headline figure. Run by material, only artificial pictures showed bias, and words showed none. Their p-curve analysis found no evidence of p-hacking.

What the authors themselves recommend

You don't have to take my reading of the table. Brunmair and Richter wrote the practical implications themselves:

We are reluctant to recommend interleaving for learning materials such as mathematical tasks, expository texts, grammar rules, foreign languages or words (i.e., category names). When interleaving is used with these materials, it could even impede learning compared to a blocked presentation of items.

Brunmair & Richter, 2019, accepted manuscript, p. 36

"Foreign languages" is in that list, in the paper cited as the evidence base for interleaving your foreign language study. They name art history, biology, medicine and geology as places to use it. They will not recommend it for the thing you're doing.

The label "words" hides how close those 13 studies sit to ordinary vocabulary study. The paper defines the category as "names that belonged to different conceptual categories, pronunciation rules, or translations in different languages." One of the three it names is Carpenter and Mueller, 2013: native English speakers learning French pronunciation rules, blocked or interleaved. Four experiments. Blocking won all four, at 4 words per rule and at 15 words per rule.

The third named study used actual flashcards. Hausman and Kornell, 2014, titled it "Mixing topics while studying does not enhance learning." Participants studied anatomy terms and Indonesian translations, either alternating between topics or doing one then the other. Four experiments, and "Mixing did not have reliable effects." Kornell co-authored the paintings study that made interleaving famous, which makes the null worth more coming from him.

I read both of those at abstract level only, from PubMed and OpenAlex. Their full texts sit behind publisher blocks I could not get past.

The evidence against my own argument

Two findings cut the other way, and leaving them out would make this page worse than the one it's correcting.

Pan and colleagues, 2019, in the Journal of Educational Psychology, taught adults to conjugate Spanish preterite and imperfect verbs. Across two weekly sessions interleaving won clearly: on the Experiment 4 delayed test, accuracy was 0.30 blocked against 0.49 interleaved, d = 0.79. They call it "the first demonstration of an interleaving effect for foreign language learning," striking to publish in 2019, given how long the advice had circulated.

Two details matter for a vocabulary deck. It's grammar, so the category being discriminated is verb tense and the learner is applying a rule. And in their single-session experiments interleaving lost or drew: Experiment 1 gave "numerically higher performance in the blocked group," Experiment 2 came out equivalent. The win needed multiple sessions.

And Pan's own discussion says this, in the middle of announcing their positive result:

language is not a single capacity but a collection of many skill subdomains, and interleaving's benefits are likely to vary by task type; for instance, it has not shown a benefit for learning vocabulary.

Pan, Tajran, Lovelett, Osuna & Rickard, 2019, General Discussion

The second finding is newer, and it's vocabulary, and it goes against me. Libersky and colleagues, 2025, in Second Language Research, taught English-speaking adults words in German or Polish, interleaving or blocking the two languages, and found an interleaving benefit in Experiment 1. Adding a break midway in Experiment 2 made the conditions perform the same, which led them to read the effect as spacing. What they interleaved was two entire languages, a coarser contrast than shuffling food words with travel words inside one deck. I'm working from the abstract: the full text is paywalled at SAGE and returned 403.

Put together, the picture I'd defend: interleaving is well evidenced for visual category learning, evidenced for L2 grammar across multiple sessions, and evidenced against for word learning. The claim that lost is the broad one.

The post I got wrong

In July 2026 I published a post recommending interleaving for vocabulary. It's still up. It cites this exact meta-analysis, and here is the sentence I wrote:

Brunmair and Richter's 2019 meta-analysis across 238 comparisons confirmed the advantage holds broadly, strongest exactly where materials are similar enough to confuse.

My own post, /blog/interleaving, published 5 July 2026

The second half is right, straight from the meta-regression: similarity between categories does raise the effect. The first half is wrong. "The advantage holds broadly" is the claim the paper was written to test and declined to support.

I had the abstract when I wrote that. The words sentence is in the abstract. I read the number I was looking for and stopped, which is the ordinary way this happens, and it's why the citation chains here point at a paper more careful than any of the pages citing it.

There's a correction note at the top of that post now, pointing here. It stays published, because deleting it hides the error rather than fixes it, and because its discrimination-learning material still holds. What doesn't hold is applying it to a vocabulary deck.

What I'd actually do with a deck

The practical answer is smaller than either post makes it sound. The words result is about how you order items while studying new material. It says nothing about spacing, a separate and far better evidenced effect, and nothing about retrieval practice. Anything a scheduler does across days is untouched by this.

So: when you meet a batch of new words, grouping them by theme for that first pass is defensible. When they come back days later in whatever order they're due, that's spacing doing its work.

Two studies point at that ordering directly. Sorensen and Woltz ran a group that transitioned incrementally from blocked into interleaved, and learners whose initial exposure was blocked did better on both implicit and explicit tests. Hwang, 2025, in Language Learning, ran the same question on 107 Korean adolescents learning English vocabulary: blocking, interleaving, or a hybrid of both. Interleaving alone posed what he calls undesirable difficulty, early blocked practice "facilitated the development of new declarative knowledge," and the hybrid held up best on long-term retention. Block the introduction, mix the review.

The uncertain middle is contrast pairs. Ser against estar, por against para. Those sit closer to rule discrimination, they're the case Pan's grammar result covers, and the between-category similarity finding predicts interleaving should help. I think mixing them is right. I can't prove it from the words row.

Where this touches what I build: Verbamor schedules cards by due date, so review order comes out mixed whatever anyone prefers. That's spacing, and the spacing evidence is strong. The claim I'm dropping is that the shuffling itself teaches you vocabulary. Two effects were riding in one sentence in my old post, and only one of them has the numbers.

If you take one thing from this page: check the moderator table before you take an effect size. The average across materials you don't study is not about you.


Sources

The meta-analysis

  • Brunmair, M., & Richter, T. (2019). Similarity matters: A meta-analysis of interleaved learning and its moderators. Psychological Bulletin, 145(11), 1029-1052. Read in full from the accepted manuscript hosted by the University of Würzburg. Table 2, p. 54: overall k = 238, g = 0.42 [0.34, 0.50]; paintings k = 70, g = 0.67; naturalistic photographs k = 19, g = 0.35; artificial pictures k = 63, g = 0.31; expository texts k = 25, g = 0.21, p = .119; mathematical tasks k = 45, g = 0.34; words k = 13, g = -0.39, p = .005, [-0.64, -0.14], I² = 18.1% in Table 2 and 18.3% in the body text, a discrepancy in the paper itself; tastes k = 3, g = 0.57, p = .24. Trim-and-fill on the total sample, p. 31: g = 0.29 [0.20, 0.38] with 23 studies missing, and no bias detected for words. Table 3, p. 55, Model 3, reference category paintings: words b = -0.48 (SE 0.23, p = .014), expository texts b = -0.56, mathematical tasks b = -0.43, familiarity b = 0.19 to 0.20 across models. The "reluctant to recommend" passage is on p. 36; the definition of the words category and its example studies are on p. 20.

The word studies

  • Carpenter, S. K., & Mueller, F. E. (2013). The effects of interleaving versus blocking on foreign language pronunciation learning. Memory & Cognition, 41, 671-682. Abstract read via PubMed, PMID 23322358: four experiments, blocked or interleaved French pronunciation rules, "In all experiments, blocking benefited the learning of pronunciations more than did interleaving, and this was true whether participants learned only 4 words per rule (Experiments 1-3) or 15 words per rule (Experiment 4)." Full text not obtained, Springer serves a bot challenge to automated fetches.
  • Hausman, H., & Kornell, N. (2014). Mixing topics while studying does not enhance learning. Journal of Applied Research in Memory and Cognition, 3, 153-160. The third words-category study named by Brunmair and Richter. Abstract via OpenAlex: participants "alternated on each trial between studying anatomy terms and Indonesian translations" against studying "one topic and then the other"; "Mixing did not have reliable effects when participants studied flashcards in a single day (Experiments 1 and 2) or on two different days (Experiments 3 and 4)." Closed access, full text not obtained.
  • Sorensen, L. J., & Woltz, D. J. (2016). Blocking as a friend of induction in verbal category learning. Memory & Cognition, 44, 1000-1013. Also named by Brunmair and Richter as a words-category study. Abstract read via PubMed, PMID 27115608: four groups, blocked, interleaved, or "an incremental transition from blocked to interleaved practice"; "Counter to current trends in the literature, in which interleaving alone has been facilitative of induction in some tasks, we found that participants whose initial exposure to the category exemplars involved blocked presentation performed better in both implicit and explicit tests of concept learning." Full text not obtained, same Springer challenge.

Evidence the other way

The arithmetic, mine

  • Standard errors recovered from the published 95% intervals as (upper − lower) / 3.92: 0.041 for the overall estimate, 0.128 for words. Difference 0.42 − (−0.39) = 0.81, pooled SE = √(0.041² + 0.128²) = 0.134, z = 6.05. Visual materials 70 + 19 + 63 = 152 of 238 effect sizes = 63.9%; words 13 of 238 = 5.5%. Trim-and-fill reduction (0.42 − 0.29) / 0.42 = 31%. The percentile readings are the normal CDF of each g: 66% at 0.42, 35% at −0.39. Brunmair and Richter publish none of these; the inputs are all from their Table 2.

Every load-bearing claim Verbamor makes is traced to its paper on the research page.

Spacing is the part with the numbers.

Verbamor schedules the vocabulary from your own lessons by when each card is due, which is the effect this page didn't have to argue about.

Download on the App Store