Verbamor Download the appGet the app
All posts Download on the App Store
Product

The flashcard that plays a tone where the word should be

A card plays a sentence with a soft tone standing in for one word, and your job is to say that word out loud before the answer appears. Get it out in the second after the tone and you actually know it. Wait for the reveal, and the card just told you the truth about that.

Here is a card from a Spanish deck. The screen shows a sentence with a gap in it: El trabajador ____ una caja pesada. Then the audio plays. You hear a voice read the whole sentence, and where the missing word belongs you hear a soft tone instead. Your job is to say the word before you tap.

The word is levanta. If it came out of your mouth in the second after the tone, you own it. If you sat there and waited for the answer, you do not, and the card has just told you so in a way a two-sided card never could.

What a cloze card actually is

Cloze means a sentence with something taken out of it. The name is from Wilson Taylor, who published the cloze procedure in Journalism Quarterly in 1953 as a way to measure how readable a piece of writing was. He deleted words at intervals and counted how many readers could restore. He took the word from closure, the idea from Gestalt psychology that a mind shown an incomplete pattern will move to complete it.

The gap is the whole mechanism. A two-sided card shows you levantar and asks for "to lift". You answer with a translation, which is a task you will never perform in a conversation. A cloze card shows you a sentence that is missing one piece, and the missing piece has to be the right word in the right form in the right place. Levanta, not levantar, because the sentence has a worker doing it right now.

That matters most for the parts of a language that only exist in context. Gender, conjugation, which preposition a verb takes, whether the adjective goes before or after. A word list cannot test any of those. A sentence with a hole in it tests all of them at once, because there is only one shape that fits.

Why the audio carries a tone instead of the word

Reading a gap on a screen and hearing a gap are different tasks, and the second one is closer to what happens when a person talks to you.

A written blank gives you time. You can look at the sentence, work out the tense from una caja pesada, reason your way to levanta, and feel good about it. That reasoning is worth something, but it is not what conversation asks for. Conversation gives you a stream of sound moving at its own speed, and the word either arrives or it does not.

So the recall side of the card carries a recording of the sentence with the answer replaced by a tone. You hear the rhythm, the words on both sides of the gap, the intonation of the whole line, and a beep where your word goes. The stress pattern of the sentence is intact, which means you are practicing against the real timing of the language rather than against your own reading speed.

The other face of the card plays the complete recording in the same voice, read straight through with the word spoken. The tone version is the test. The full version is the model you are copying.

Say it out loud, because the mouth is doing work

MacLeod, Gopie, Hourihan, Neary and Ozubko ran eight recognition experiments and named the result the production effect (Journal of Experimental Psychology: Learning, Memory, and Cognition, 2010, 36(3), 671-685). Producing a word aloud at study, compared with reading it silently, improved memory for it. Their explanation is distinctiveness: the spoken item gets a record with extra information attached, and at test that record is easier to pick out.

Forrin and MacLeod later split the effect into its parts (Memory, 2018, 26(4), 574-579). They compared four conditions: reading aloud yourself, hearing a recording of your own voice, hearing someone else read aloud, and reading silently. Memory fell in a gradient across those four, with hearing yourself sitting between speaking and hearing another person. Two things are doing the work, the motor act of speaking and the fact that the voice is yours.

One honest caveat travels with this. The 2010 experiments found the effect in mixed lists, where some items were spoken and others were not, within the same person. It is a distinctiveness effect, so it depends on the spoken items standing out from silent ones. A review session where you speak every single answer is not the design that was tested. Speaking still gets you the motor and self-referential components, and it forces you to commit to an answer before the flip, which is worth doing on its own terms.

The commitment problem this solves

Koriat and Bjork named the mechanism behind that gap (Journal of Experimental Psychology: Learning, Memory, and Cognition, 2005, 31(2), 187-194). They call it an illusion of competence. When you study a pair, you see the cue and the answer together. When you are tested, you see the cue alone and have to produce the answer. Judging your own learning while the answer sits in front of you rates the wrong thing. You rate how well the pair fits together, which is easy with both halves visible. What decides recall is whether the cue alone brings the answer back. Those two are different, and the first one feels like the second.

A spoken answer removes the room where that illusion lives. You cannot quietly revise "I basically had it" into a pass once the word is already out of your mouth and wrong. The card makes the commitment audible, then plays the correct version so you hear the difference immediately.

This is the same reason the image on the card is a hint and not a shortcut. Carpenter and Olson tested pictures against native-language translations for Swahili vocabulary (Journal of Experimental Psychology: Learning, Memory, and Cognition, 2012, 38(1), 92-101). Pictures did not beat translations at first, and learners were more overconfident recalling from a picture than from a translation. Pictures only helped once that overconfidence was corrected by retrieval practice. An image next to a sentence you have to complete out loud is an image that has to survive a test.

What this looks like after four weeks

Say you record one 50-minute tutor lesson a week and keep 12 words and phrases from each. That is 48 cards after four weeks. Cards are scheduled by FSRS, the same algorithm family Anki offers, though Anki ships SM-2 by default and FSRS is opt-in there. FSRS-6 fits 21 parameters to your own review history rather than applying one fixed formula to everyone, and it schedules to a desired retention you set, 0.90 by default in Anki.

Here is one card's ladder. Take card number one from lesson one and press Good every time. Running the reference FSRS-6 implementation (the fsrs Python package, version 6.3.2, at default parameters and 0.90 retention) that card comes back on day 1, again on day 3, then day 13, then day 60. Four sightings in two months, and only two of them inside your first four weeks. Diff that against your own screen.

The catch is that the ladder above assumes you never forget, and you will. Simulating all 48 cards with a forget rate matching FSRS's own retrievability estimate, the load is spiky rather than flat. Lesson day is the heavy one, around 35 to 43 reviews, because 12 new cards each get seen more than once. The other six days of week four run a median of 8 reviews, with the middle 80% of days falling between 4 and 18. Across the whole 28 days the median is about 361 reviews total.

I am not going to convert that into minutes for you, because the app does not record how long a review takes and I will not invent the number. Time one session with a stopwatch and you will have a figure that is true for you rather than one I guessed.

What changes in your speech is narrower than "you get fluent" and more useful. The cards from lesson one are the words your own tutor used with you, in sentences about the thing you were actually discussing, so when that topic comes back you have the phrase in the mouth rather than in the notebook. The forms come with it, because you never practiced levantar alone.

Who should skip this

If you are in your first two weeks of a language, a sentence with a gap is the wrong tool. You need enough grammar to know what shape fits the hole, which in practice means you can already conjugate a regular present tense and recognise the common gendered articles. Elizabeth Bjork and Robert Bjork make the point directly in "Making Things Hard on Yourself, But in a Good Way", their chapter in Psychology and the Real World (Worth Publishers, 2011, pages 56-64): difficulties help because they trigger encoding and retrieval processes that support learning, but if the learner lacks the background knowledge to respond successfully, they become undesirable difficulties. Do a beginner course first and come back.

Skip it too if you never speak the language out loud with anyone. Recall of a spoken sentence under time pressure is training for a specific situation. If reading is your goal, a reading-focused deck costs you less.

And if you already run Anki with a cloze template, audio on your cards, and the patience to cut clips yourself, you have most of this. What you are buying elsewhere is the part where the sentence, the two recordings and the image get made for you from a lesson you recorded, instead of you making them at midnight.

Record your next lesson, keep ten words from it, and on the first review say every answer out loud before you tap. Count how many arrived before the tone ended. That number is the only one worth watching.

Say the word before you see it.

Verbamor builds every card from your own lessons with a tone standing in for the answer, so review means producing the word out loud, the same thing a real conversation actually demands.

Download on the App Store