How to grade yourself honestly on a flashcard
To grade yourself honestly on a flashcard, say the word before you flip, and if the answer arrived after the flip, grade it Again, with no exceptions for "I basically had it." An Again costs one extra review tomorrow; a fake Good costs 3 weeks of silent decay, because the scheduler believes whatever grade you give it.
The card shows "the forest" and a blank. You think... bosque? Maybe busque? Something with a b. You flip it: bosque. Close enough, you tap Good, and the word sails away for 3 weeks.
Except you didn't know bosque. You recognized it after the fact, which is a different thing entirely. Now a word you half-know is scheduled like a word you own.
Three weeks from now it'll be gone, and it'll take extra reviews to drag it back. The whole transaction cost you one moment of honesty.
01 / 06Inflate the grades and the intervals inflate with them
A modern scheduler like FSRS, an open-source scheduling algorithm, is a prediction machine. It watches your grades, fits a memory model per card, and times the next review for the moment you're about to forget.
Its entire view of your memory comes through one channel: the grades.
Feed FSRS inflated grades and it builds an inflated model. Intervals stretch beyond what the memory can hold. Words start failing at 3 weeks that should've been reviewed at 5 days.
And the deck develops that demoralizing pattern where everything feels simultaneously "reviewed" and shaky.
The scheduler has one channel into your memory: the grades. Lie on the input and the output lies back.
Grade honestly and FSRS is genuinely good at its job. The difference between those two decks is calibration: how well your grades match what you actually recall.
Noticing what the grade is for helps. The grade is a sensor reading, and no one's GPA is riding on it. How faded was this memory when you caught it?
The scheduler needs that reading the way a thermostat needs the thermometer. And a thermostat wired to a flattering thermometer doesn't produce a warmer honest house.
02 / 06The students who over-rated themselves retained the least
The bad news from metacognition research (the study of how we judge our own minds): your default self-assessment runs hot, and in predictable ways.
Koriat and Bjork (2005) named the core problem illusions of competence. When you study with the answer in view, you judge how well you know it while looking at it. Everything feels knowable with the answer on the table.
The judgment you need (could I produce this cold?) is precisely the one the flipped card no longer measures.
And grade inflation has a cost you can measure. Dunlosky and Rawson (2012) had students learn definitions while judging their own accuracy.
The overconfident ones ended retention tests with the worst scores, because material judged "known" gets retired from practice early. Their paper's blunt subtitle: overconfidence produces underachievement.
Add the fluency trap (fast-feeling recognition reads as knowledge, see the testing effect) and the pattern is clear. Every bias points the same direction, toward over-grading.
What helps here is a procedure.
03 / 06The same judgment, made after a delay, predicts recall far better
Before the procedure, some genuinely good news: self-assessment is a skill with a known upgrade path.
Nelson and Dunlosky (1991) studied judgments of learning, your own guess at how well you'll remember something. Made immediately after studying an item, they're mediocre predictors of later recall. The same judgments made after a delay are startlingly accurate.
The immediate judgment reads the afterglow. The delayed one has to actually probe the memory, which is the honest test.
Your review session exploits the delayed-judgment effect automatically. Every card arrives days after you last saw it. So the judgment you're making at the moment of recall is the accurate kind, provided you make it before the flip.
The commit-first habit (say it, then look) is a delayed judgment of learning wearing work clothes.
Calibration also improves with feedback. Learners who grade strictly for a few weeks start noticing which internal signals actually predict success.
The word arrives whole, or it assembles itself letter by letter. Confidence has a shape to it, or it's vague warmth.
That noticing transfers, usefully, to real conversation. There, knowing whether you actually know a word decides whether you start the sentence.
04 / 06When should you grade a card Again?
Grade a card Again whenever the word didn't come out of your head before the flip. Again is the first of Verbamor's 4 grades. Here's the honest mapping, and the only hard rule is the first one:
The bright line: if the answer arrived after the flip, it's an Again. No exceptions for "I basically had it." Basically-had-it is the exact state the next review exists to fix, and it can only fix it if the scheduler knows.
One more habit makes the rubric self-enforcing: commit before you flip. Say the word out loud, or type it, before the answer appears. A spoken guess can't be quietly revised into a pass after the fact.
This is half the reason Verbamor cards make you produce into a blank rather than flip-and-judge.
05 / 06Rule strict on the edge cases, because the bias never stops leaning
Real reviews produce edge cases. My rulings, for what they're worth:
- Right word, wrong gender. La bosque? Hard, at best. The article is part of the word; you'll say it wrong at dinner exactly the way you said it wrong on the card.
- Right word, mangled pronunciation. If you'd be understood, Hard. If your tutor would squint, Again.
- Took 10 seconds. Still counts, that's what retrieval looks like near the edge. Hard or Good depending on how ugly the search felt.
- Got the word from the image, not the sentence. Fine. The image is a legitimate route (that's dual coding doing its job). Good.
- Synonym instead. You said selva for forest? You know a word; you don't know this card. Again, and no shame in it.
You'll notice the rulings trend strict. That's the point of having rulings: to lean against a bias that never stops leaning the other way.
06 / 06The Again button costs seconds, and a fake pass costs three weeks
The psychology underneath all of this is that Again feels like losing. In fact an Again is one extra review tomorrow, maybe 8 seconds of your life.
And the scheduler quietly rebuilds the word on a schedule that matches reality.
A fake Good costs more: 3 weeks of silent decay, a failed retrieval later anyway, and a model of the word that's now wrong in both directions.
Kornell's work on retrieval failure even shows the miss itself has value. A failed attempt followed by the answer still strengthens the memory more than a smooth flip ever did.
So the honest grade is also, conveniently, the selfish one. Tap Again like it's nothing, because it is.
The learners whose decks hold up at month 12 are the ones whose grades meant something.
And if you've been generous for months, don't rebuild the deck. Just start grading straight today.
The scheduler re-fits to your real performance within a few weeks of clean data, the way any model recovers once the sensor stops lying.
Sources
The scheduler reads your grades
- Ye, J., Su, J., & Cao, Y. (2022). A stochastic shortest path algorithm for optimizing spaced repetition scheduling. Proceedings of KDD '22.
Why we over-grade ourselves
- Koriat, A., & Bjork, R. A. (2005). Illusions of competence in monitoring one's knowledge during study. Journal of Experimental Psychology: Learning, Memory, and Cognition, 31(2), 187–194.
- Dunlosky, J., & Rawson, K. A. (2012). Overconfidence produces underachievement: Inaccurate self evaluations undermine students' learning and retention. Learning and Instruction, 22(4), 271–280.
- Rhodes, M. G., & Castel, A. D. (2008). Memory predictions are influenced by perceptual information: Evidence for metacognitive illusions. Journal of Experimental Psychology: General, 137(4), 615–625.
Calibration is a skill
- Nelson, T. O., & Dunlosky, J. (1991). When people's judgments of learning (JOLs) are extremely accurate at predicting subsequent recall: The "delayed-JOL effect." Psychological Science, 2(4), 267–270.
Failed tries still help
- Kornell, N., Hays, M. J., & Bjork, R. A. (2009). Unsuccessful retrieval attempts enhance subsequent learning. Journal of Experimental Psychology: Learning, Memory, and Cognition, 35(4), 989–998.
Grades in, memory out.
Verbamor's scheduler learns your memory from every grade you give it, and repays honesty with intervals that actually hold.
Start 7 days free