Verbamor Download the appGet the app
All posts Download on the App Store
Memory science

The input hypothesis still cannot be tested

Krashen never said how to measure i+1, so no result can contradict him. A 2025 meta-analysis of 82 reading interventions found the biggest effects where learners had less choice and were held to account, which is the opposite of what the doctrine prescribes.

Stephen Krashen's input hypothesis is the reason your teacher tells you to watch more television and stop drilling verbs. It's the spine of comprehensible input, of free voluntary reading, of every "just get more input" reply in every language forum.

It has had one structural problem since 1984. You cannot run an experiment that would prove it wrong.

The hypothesis with no ruler

The claim is that you acquire language when you understand input containing, in Krashen's words, "structure a bit beyond our current level of competence (i + 1)".

To test that you need to measure two things: a learner's level, i, and the size of the step, +1. Krashen never specified how to do either.

"Although the IH has been criticized for being untestable (Krashen has not clarified how to operationalize i + 1), many theories since have acknowledged the fundamental role of input in language learning."

Sangers et al., 2025, in the introduction to their own meta-analysis. Their original sets the middle clause off with dashes, replaced here with parentheses.

Without a ruler for i+1, every outcome fits. Learner improved? The input contained i+1. Learner didn't? Then it didn't, or the affective filter was up. No observation comes back and says the hypothesis is wrong, which is what Kevin Gregg meant in 1984 when he wrote that each of Krashen's five hypotheses is marked by "undefined or ill-defined terms, unmotivated constructs, lack of empirical content and thus of falsifiability, lack of explanatory power".

That sentence is 42 years old, and the 2025 paper quoted above restates the gap as settled background before getting on with its work.

Being untestable isn't the same as being wrong, and I want to be careful here. Input obviously matters. Nobody acquired a language without it. The problem is narrower: a claim built this way absorbs any result, so evidence can't improve it. It sits there collecting agreement.

The field landed on a split verdict, which is why that quoted sentence turns at the comma. Behind its second half is Lichtman and VanPatten's 2021 reassessment, "Was Krashen right? Forty Years Later". I have not read it, it's paywalled with no repository copy, so treat it as a pointer to the best case against my reading. Meanwhile a more answerable question sits next door. Which ingredients actually move the numbers?

What 82 reading interventions did

Sangers, van der Sande, Welie, Dobber and van Steensel pooled 73 studies covering 82 extensive-reading interventions. Extensive reading is the classroom form of the doctrine: lots of easy text, self-selected, read for pleasure, no tests attached.

It works. The immediate effect across all 82 interventions is d = 0.38, standard error 0.07. By domain: general language proficiency 0.53, vocabulary 0.50, oral proficiency 0.42, decoding and fluency 0.38, motivation 0.34, reading comprehension 0.31, writing 0.31. All positive, all significant, small to medium. That deserves to be the headline and the authors treat it as one.

Then look at which interventions produced it.

0
d = 0.73 when learners' text choice was limited
0
d = 0.22 when choice was boundless, the doctrine's own prescription
0
d = 0.01 when nothing held the learner to account
Sangers et al. 2025, Table 3. Moderator contrasts across 82 interventions.

Text limitation, meaning learners were steered toward books at their reading level: d = 0.73, 95% CI [0.48, 0.99], k = 25. Boundless choice: d = 0.22, CI [0.04, 0.40], k = 57. Q = 10.52, p < .01.

Accountability, meaning a reading log or a short quiz: d = 0.51, CI [0.35, 0.68], k = 59. Without it, d = 0.01, CI [−0.28, 0.30], k = 23. Q = 8.72, p < .01.

Sit with the second one. Extensive reading with nothing checking that you read is 0.01, with an interval spanning zero in both directions. On this evidence it's indistinguishable from doing nothing.

Now put those next to the doctrine's own rulebook. Day and Bamford's 10 principles of extensive reading, which the paper quotes as the regularly cited standard, include number 3, "learners choose what they want to read", and number 6, "reading is its own reward".

Those are the 2 the moderators contradict. Principle 3 is the free-choice condition at d = 0.22, and principle 6 rules out the accountability separating 0.51 from 0.01. Behind the choice principle is Krashen's affective filter, where autonomy lowers anxiety and lets input through. Here the autonomous condition is the weaker one.

Those 2 are the only significant moderators in their Table 3, which tests 12 intervention characteristics. Here are the other 10, all of them: graded readers Q = 0.000, digital support 0.001, rewards 0.004, home books 0.03, interaction 0.05, home activities 0.10, text selection 0.13, teacher help 0.23, text medium 1.10, researcher-run 1.11.

So 10 of the 12 things you could change about a reading programme did nothing measurable, and the 2 that moved it are the 2 the doctrine argues against.

The cross-tab the paper doesn't print

The moderators are reported one at a time, leaving an obvious question open: how many of these 82 interventions were the pure article, free choice with no accountability?

Their appendix codes every intervention on every characteristic, so I built the 2×2 myself from Table 6. It reproduces their published marginals exactly, 25 text-limited and 59 accountable, which checks the parse.

17 of 82 interventions were free-choice with no accountability, or 20.7 percent. The 82 split 17 / 40 / 6 / 19 across the four cells.

My cross-tabulation of Sangers et al.'s Appendix Table 6

The largest cell, 40 interventions, is free choice plus accountability. Nearly half the evidence base for "extensive reading works" comes from programmes that let you pick your book and then asked you to prove you'd read it. 19 more were limited and accountable, and just 6 were limited without.

So the literature cited for free voluntary reading is mostly testing self-selection with a teacher checking, which is a different intervention with a different name. When a forum reply cites "the extensive reading research" to argue against structure, the research it points at mostly had structure in it.

The 2×2 settles something the separate tables can't. Both moderators cover the same 82 interventions, which raises the worry that they're one finding counted twice: maybe the programmes limiting text choice were the same ones adding quizzes.

They're close to independent. 70.2 percent of free-choice interventions had accountability, against 76.0 percent of text-limited ones: an odds ratio of 1.35, a phi coefficient of 0.06. So the accountability result is no shadow of the text-limitation result, and the 2 are worth adding separately.

Two caveats on my arithmetic. The 2×2 counts interventions, unweighted by precision, so it describes what the literature studied rather than what it found. A per-cell effect size would settle much more, but the appendices publish characteristics and study quality without a per-intervention d, so those cell means can't be recovered. The paper reports no total participant N either.

The part that weakens my own argument

Two things work against the reading I've just given you, and a page that hid them would deserve less trust.

First, the d = 0.01 is softer than it looks. The authors coded missing information as absence. Their rule, verbatim: "In case a study did not provide information on a characteristic, we assumed that it was not part of the intervention (thus, a 'no' was coded)."

So those 23 interventions mix programmes that genuinely had no accountability with programmes whose write-ups never mentioned any. Reporting here is thin: the same paper notes 45 of 73 studies reported nothing on implementation fidelity. The 0.01 is a floor estimate over a contaminated cell, so treat the direction as real and the magnitude as unreliable.

The same rule cuts the other way on the finding I lean on hardest. Text limitation used the identical convention, so the 57-intervention free-choice cell also holds programmes that limited text choice and never wrote it down, each misfiled into the d = 0.22 group. Miscoding like that pulls the 2 cells together, so the gap between 0.73 and 0.22 is more likely understated than exaggerated. That's the direction of the bias, not a correction for it.

Second, this is my reading, not theirs. The authors present none of this as a refutation of Krashen. They're warm about extensive reading throughout, call it "underutilized", want more of it in schools, and cite Krashen's own 2011 meta-analysis approvingly on the way past.

Their conclusion is the strongest version of the case against me, so here it is whole. The analysis "supports the assumption that encouraging second/foreign language learners to engage in sustained, independent, and self-selected reading promotes a range of second/foreign language skills. This is particularly true when ingredients are included that allow a suitable match of reader and text, and that increase accountability."

Note the word self-selected in their own summary. They read the moderators as ingredients that improve self-selected reading. I read them as evidence that self-selection carries less of the load than the doctrine claims. Both readings fit the table, and mine is an argument rather than a finding they endorse.

And d = 0.38 sits inside the range other meta-analyses keep landing in. Nakanishi's 2015 analysis of 34 comparisons gave d = 0.46 against control groups and 0.71 pre-post. Everyone agrees reading works. The open argument is which ingredient does the work.

What to do with a theory you can't test

Keep the reading. Drop the purity.

Both moderators that helped are structure bolted onto the reading, and structure is what the doctrine treats as contamination. Pick books near your level instead of whatever looks interesting, and put something at the end that makes you account for what you read.

That second one is cheap. A reading log is a line per chapter, and a quiz can be your tutor asking what happened. The 59 interventions with something like this pooled to d = 0.51 against 0.01 for the 23 without.

The level advice has independent support: the coverage arithmetic behind your first novel says the same from the vocabulary side, which is why graded readers earn their place for the opposite reason they're sold. On what reading alone deposits, the incidental learning rates are the honest picture.

This is where Verbamor sits, so I'll be plain about the interest. It takes vocabulary from your own lessons and schedules it as retrieval practice, one form of the accountability the 59 had. It gives you no input and won't replace the reading. For that half, a library card and a graded reader series beat any subscription.

The honest position on Krashen is that he was directionally right about input and gave the field a claim it can't check. 42 years on, the moderator tables are where the progress is, because their numbers can come back and disagree with you.


Sources

The 2025 meta-analysis

  • Sangers, N. L., van der Sande, L., Welie, C., Dobber, M., & van Steensel, R. (2025). Learning a Language Through Reading: A Meta-analysis. Educational Psychology Review, 37(4), Article 96. CC-BY. 73 studies, 82 interventions. Immediate effects k = 82, d = 0.38, SE = 0.07. Table 3: text limitation yes k = 25 d = 0.73 [0.48, 0.99] vs no k = 57 d = 0.22 [0.04, 0.40], Q = 10.52, p < .01; accountability yes k = 59 d = 0.51 [0.35, 0.68] vs no k = 23 d = 0.01 [−0.28, 0.30], Q = 8.72, p < .01. Coding rule, p. 11: "In case a study did not provide information on a characteristic, we assumed that it was not part of the intervention (thus, a 'no' was coded)." Discussion: extensive reading "seems particularly beneficial when book selection is made easier and when there is an external impetus for reading". Conclusion: the analysis "supports the assumption that encouraging second/foreign language learners to engage in sustained, independent, and self-selected reading promotes a range of second/foreign language skills. This is particularly true when ingredients are included that allow a suitable match of reader and text, and that increase accountability." Table 3 tests 12 intervention characteristics; text limitation and accountability are the only significant moderators. 45 of 73 studies did not report implementation fidelity. The paper reports no total participant N. Full text read 3 September 2026.

The falsifiability critique

  • Krashen's formulation, quoted in Sangers et al. from Krashen, S. (1992), "The input hypothesis: An update", p. 21: acquisition occurs when input contains "structure a bit beyond our current level of competence (i + 1)".
  • Lichtman, K., & VanPatten, B. (2021). "Was Krashen right? Forty Years Later", Foreign Language Annals, 54(2), 283-305, doi:10.1111/flan.12552. Not read. OpenAlex records it as closed access with no repository full text. Named here only because Sangers et al. cite it as the counterweight to the untestability complaint.
  • Day, R., & Bamford, J. (2002), the 10 principles of extensive reading, quoted here as listed in Sangers et al. p. 2: principle 3 "learners choose what they want to read" and principle 6 "reading is its own reward". Cited at second hand, from the Sangers text.
  • Gregg, K. R. (1984). "Krashen's Monitor and Occam's Razor", Applied Linguistics, 5(2), 79-100, doi:10.1093/applin/5.2.79. Not read in the original. OpenAlex records it as closed access with no repository copy, and it was not retrieved. The quoted sentence ("undefined or ill-defined terms, unmotivated constructs, lack of empirical content and thus of falsifiability, lack of explanatory power", attributed to p. 94) is taken from Lai, W., & Wei, L. (2019), "A Critical Evaluation of Krashen's Monitor Model", Theory and Practice in Language Studies, 9(11), which was read in full.

My arithmetic

  • The 2×2 cross-tabulation of text limitation against accountability was built by parsing all 82 rows of Appendix Table 6 in the Sangers PDF. Cells: free choice with no accountability 17, free choice with accountability 40, limited with no accountability 6, limited with accountability 19. Sum 82. Column totals 25 text-limited and 59 accountable, which match Table 3 exactly and validate the parse. 17 / 82 = 20.7 percent. This cross-tabulation is not printed in the paper. It counts interventions, unweighted by effect-size precision.

Comparison figures

  • Nakanishi, T. (2015), as reported in Sangers et al.: 34 pretest-posttest comparisons including 18 experimental vs. control studies, d = 0.71 pre-post and d = 0.46 against control groups. Cited here at second hand, from the Sangers text.

Every load-bearing claim Verbamor makes is traced to its paper on the research page.

Accountability, without the reading log.

Verbamor turns the words from your own lessons into scheduled retrievals, so something checks that the input landed.

Download on the App Store