Verbamor Start 7 days free
All posts Start 7 days free

How to grade yourself honestly on a flashcard

To grade yourself honestly on a flashcard, say the word before you flip, and if the answer arrived after the flip, grade it Again, with no exceptions for "I basically had it." An Again costs one extra review tomorrow; a fake Good costs 3 weeks of silent decay, because the scheduler believes whatever grade you give it.

The card shows "the forest" and a blank. You think... bosque? Maybe busque? Something with a b. You flip it: bosque. Close enough, you tap Good, and the word sails away for 3 weeks.

Except you didn't know bosque. You recognized it after the fact, which is a different thing entirely. Now a word you half-know is scheduled like a word you own.

Three weeks from now it'll be gone, and it'll take extra reviews to drag it back. The whole transaction cost you one moment of honesty.

01 / 06

Inflate the grades and the intervals inflate with them

A modern scheduler like FSRS, an open-source scheduling algorithm, is a prediction machine. It watches your grades, fits a memory model per card, and times the next review for the moment you're about to forget.

Its entire view of your memory comes through one channel: the grades.

Feed FSRS inflated grades and it builds an inflated model. Intervals stretch beyond what the memory can hold. Words start failing at 3 weeks that should've been reviewed at 5 days.

And the deck develops that demoralizing pattern where everything feels simultaneously "reviewed" and shaky.

The scheduler has one channel into your memory: the grades. Lie on the input and the output lies back.

Grade honestly and FSRS is genuinely good at its job. The difference between those two decks is calibration: how well your grades match what you actually recall.

Noticing what the grade is for helps. The grade is a sensor reading, and no one's GPA is riding on it. How faded was this memory when you caught it?

The scheduler needs that reading the way a thermostat needs the thermometer. And a thermostat wired to a flattering thermometer doesn't produce a warmer honest house.

02 / 06

The students who over-rated themselves retained the least

The bad news from metacognition research (the study of how we judge our own minds): your default self-assessment runs hot, and in predictable ways.

Koriat and Bjork (2005) named the core problem illusions of competence. When you study with the answer in view, you judge how well you know it while looking at it. Everything feels knowable with the answer on the table.

The judgment you need (could I produce this cold?) is precisely the one the flipped card no longer measures.

And grade inflation has a cost you can measure. Dunlosky and Rawson (2012) had students learn definitions while judging their own accuracy.

The overconfident ones ended retention tests with the worst scores, because material judged "known" gets retired from practice early. Their paper's blunt subtitle: overconfidence produces underachievement.

Add the fluency trap (fast-feeling recognition reads as knowledge, see the testing effect) and the pattern is clear. Every bias points the same direction, toward over-grading.

What helps here is a procedure.

03 / 06

The same judgment, made after a delay, predicts recall far better

Before the procedure, some genuinely good news: self-assessment is a skill with a known upgrade path.

Nelson and Dunlosky (1991) studied judgments of learning, your own guess at how well you'll remember something. Made immediately after studying an item, they're mediocre predictors of later recall. The same judgments made after a delay are startlingly accurate.

The immediate judgment reads the afterglow. The delayed one has to actually probe the memory, which is the honest test.

Your review session exploits the delayed-judgment effect automatically. Every card arrives days after you last saw it. So the judgment you're making at the moment of recall is the accurate kind, provided you make it before the flip.

The commit-first habit (say it, then look) is a delayed judgment of learning wearing work clothes.

Calibration also improves with feedback. Learners who grade strictly for a few weeks start noticing which internal signals actually predict success.

The word arrives whole, or it assembles itself letter by letter. Confidence has a shape to it, or it's vague warmth.

That noticing transfers, usefully, to real conversation. There, knowing whether you actually know a word decides whether you start the sentence.

04 / 06

When should you grade a card Again?

Grade a card Again whenever the word didn't come out of your head before the flip. Again is the first of Verbamor's 4 grades. Here's the honest mapping, and the only hard rule is the first one:

The line that decides a grade is the flip, not the feeling Verbamor’s four grades, honestly applied. Only the first one is a rule; the other three are a scale for how ugly the search was.
Again didn't produce it Blank, wrong word, or "knew it once I saw it." Recognition after the flip counts as a miss.
Hard produced it, ugly Long struggle, wrong gender first, or needed the image hint to surface it.
Good produced it, solid A beat of effort, then the right word, right form. This is the everyday grade.
Easy instant, no search The word was just there, like an English word would be. Rare by design.
The line that matters is Again vs. everything else: did the word come out of your head before the flip?
Product guidance, not a measured result: Verbamor’s rubric. On why post-flip judgment fails, Koriat & Bjork 2005.

The bright line: if the answer arrived after the flip, it's an Again. No exceptions for "I basically had it." Basically-had-it is the exact state the next review exists to fix, and it can only fix it if the scheduler knows.

One more habit makes the rubric self-enforcing: commit before you flip. Say the word out loud, or type it, before the answer appears. A spoken guess can't be quietly revised into a pass after the fact.

This is half the reason Verbamor cards make you produce into a blank rather than flip-and-judge.

05 / 06

Rule strict on the edge cases, because the bias never stops leaning

Real reviews produce edge cases. My rulings, for what they're worth:

  • Right word, wrong gender. La bosque? Hard, at best. The article is part of the word; you'll say it wrong at dinner exactly the way you said it wrong on the card.
  • Right word, mangled pronunciation. If you'd be understood, Hard. If your tutor would squint, Again.
  • Took 10 seconds. Still counts, that's what retrieval looks like near the edge. Hard or Good depending on how ugly the search felt.
  • Got the word from the image, not the sentence. Fine. The image is a legitimate route (that's dual coding doing its job). Good.
  • Synonym instead. You said selva for forest? You know a word; you don't know this card. Again, and no shame in it.

You'll notice the rulings trend strict. That's the point of having rulings: to lean against a bias that never stops leaning the other way.

06 / 06

The Again button costs seconds, and a fake pass costs three weeks

The psychology underneath all of this is that Again feels like losing. In fact an Again is one extra review tomorrow, maybe 8 seconds of your life.

And the scheduler quietly rebuilds the word on a schedule that matches reality.

A fake Good costs more: 3 weeks of silent decay, a failed retrieval later anyway, and a model of the word that's now wrong in both directions.

Kornell's work on retrieval failure even shows the miss itself has value. A failed attempt followed by the answer still strengthens the memory more than a smooth flip ever did.

So the honest grade is also, conveniently, the selfish one. Tap Again like it's nothing, because it is.

The learners whose decks hold up at month 12 are the ones whose grades meant something.

And if you've been generous for months, don't rebuild the deck. Just start grading straight today.

The scheduler re-fits to your real performance within a few weeks of clean data, the way any model recovers once the sensor stops lying.


Sources

The scheduler reads your grades

Why we over-grade ourselves

Calibration is a skill

Failed tries still help

Grades in, memory out.

Verbamor's scheduler learns your memory from every grade you give it, and repays honesty with intervals that actually hold.

Start 7 days free
Verbamor screenshot: a one-minute French recording, transcribed, with the words the learner reached for listed underneath (du coup, le quai, voilà) and ticked to add to the deck. Finished cards in Spanish, Portuguese and Italian fan out beside the phone.