Verbamor Download the appGet the app
Verbamor Download the app

Tools that review vocabulary from your lesson recordings

Going from an audio file to a word you get tested on in 3 weeks takes 4 separate stages. Most tools do exactly one.

People ask for "a tool that reviews vocabulary from my lesson recordings" as though it's one product category. It's 4 jobs, and the reason the search is frustrating is that most software does 1 of them well and hands you the rest.

Here's the chain, what each link costs, and 3 combinations that actually work end to end.

The 4 stages

Audio goes in. A scheduled review comes out. In between:

  1. Transcription. Audio to text, ideally knowing who spoke.
  2. Selection. Deciding which of the 400 words in that transcript are worth your time.
  3. Card construction. Turning a chosen word into a question with a sentence, a blank, audio.
  4. Scheduling. Deciding what you see today.

Skip stage 4 and you have a document. Skip stage 2 and you have 400 cards you'll abandon by Sunday.

Stage 1: transcription

This stage is basically solved, and it's the one people worry about most.

General-purpose transcription handles a language lesson fine. What matters more than raw accuracy is speaker labels. A lesson transcript where your tutor's fluent Spanish and your halting Spanish are one undifferentiated block is much less useful, because the words worth keeping are usually the ones your tutor said and you didn't.

The second thing that matters is length. Lessons run 45 to 90 minutes, and plenty of tools are built around short clips. Check the ceiling before you rely on it.

A caution on cost: some transcription services bill the whole file regardless of what you need from it. If you only want the last 20 minutes, trim before uploading rather than after.

Stage 2: picking the words

This is the stage that decides whether the whole thing works, and it's the one most stacks skip.

A 60-minute lesson transcript holds 6,000 to 9,000 words. Maybe 30 are worth keeping. Perhaps 9 are worth keeping this week.

Frequency filtering alone gets you the wrong answer, because it surfaces words that are common in general Spanish rather than words that are new to you. The useful signal is different: the word your tutor supplied when you stalled, the correction to a mistake you keep making, the phrase carrying grammar you haven't automated.

Doing this by hand means scrubbing the recording, which is slow, and you'll do it twice before you stop. Doing it by rule means a lot of cards for words you already knew.

Stage 3: building the card

A word on its own is a bad card. The sentence it appeared in, with the word blanked out, is a good one, because it makes you produce the word rather than recognise it.

Add the article for nouns, so gender is part of what you produce. Add audio, since the written form quietly bends your pronunciation: Bassetti and Atkinson found in 2015 that spelling was still distorting pronunciation in learners years into instruction. Add an image where the word is concrete.

Do it manually and this is 90 seconds a card, which is the tax that kills most systems around week 3.

Stage 4: scheduling

Any spaced repetition scheduler beats no scheduler by an enormous margin. Cepeda and colleagues reviewed 317 experiments in 2006 and spacing wins consistently.

Between modern schedulers the gap is much smaller than the internet suggests. FSRS tracks difficulty, stability and retrievability per card and fits its parameters to your own review history; SM-2 is a fixed formula from 1987. FSRS is better. It's not the difference between working and not working, and anyone telling you the scheduler is the main event is selling one.

Three stacks that work

The manual stack. Any recorder, your own ears, Anki, and the manual method for the card writing itself. Costs $0 (or $24.99 once on iOS) and about 25 minutes per lesson. Pedagogically the strongest, because you do stages 2 and 3 yourself and that effort builds memory before the first review. Slamecka and Graf called it the generation effect in 1978. It fails on discipline, not on design.

The assembled stack. A transcription service, then a language model to pull candidate words, then a CSV into Anki. Flexible, cheap per lesson, and you own every piece. The friction is that it's 4 manual steps every time, and the CSV import is where people quietly give up.

The single-app stack. Something that owns all 4 stages. Verbamor is the one I build, so discount accordingly: record or import up to 90 minutes, transcription with speaker labels, the transcript mined for words and phrases, you tick what you want, cards built with generated audio and an image, scheduled with FSRS. The trade is that you're not choosing each stage's best tool, and you've given up the generation effect at stage 3.

Where these break

Three failure modes, all of them common.

No stage 2. The stack transcribes beautifully and generates 200 cards. You do 200 the first day, 40 the next, none after that. Selection is the difference between a deck and a landfill.

A handoff you have to do by hand. Every manual step between stages is a place to stop, and the CSV export is the classic one. It works great in week 1.

The trade automation makes at stage 3 is worth understanding before you pick a stack: what it costs you is a real thing rather than a caveat.

Recording with no plan. The most common of all. Recording feels productive and costs nothing, so people accumulate 30 hours of audio they'll never open. A recording you don't process is worth less than notes you do.

Questions people actually ask

What tools review vocabulary from a lesson recording?

No single category covers it, because it's 4 jobs: transcription, choosing which words matter, building the cards, and scheduling the reviews. Some stacks chain a transcription service to a language model to Anki. Some apps, Verbamor included, do all 4 in one place. Transcription alone gets you a document, not a review.

Can I just transcribe my lesson and make cards from the transcript?

You can, and the transcription part is easy now. The hard part is selection: a 60-minute lesson is 6,000 to 9,000 words and maybe 30 are worth keeping. Making cards from everything produces a deck you abandon within a week, so the filtering step is what determines whether it works.

Does the transcript need speaker labels?

It helps a lot. The words worth keeping are usually the ones your tutor produced and you didn't, so a transcript that separates their speech from yours makes the selection step much easier. Without labels you're reading an undifferentiated block and guessing who said what.

How long can a lesson recording be?

Depends on the tool, and it's worth checking before you rely on one. Lessons commonly run 45 to 90 minutes while plenty of transcription tools are built for short clips. Some services also bill for the whole file no matter which part you need, so trim before uploading rather than after.

Is FSRS worth switching for?

The big win is spacing itself. Cepeda and colleagues reviewed 317 experiments in 2006 and distributed practice beats massed practice consistently. FSRS improves on SM-2 by fitting parameters to your own review history rather than using a fixed 1987 formula, so it's better, though the gap between 2 modern schedulers is far smaller than the gap between scheduling and not scheduling.

Should I record every lesson?

Only if you have a plan for processing them. Recording feels productive and costs nothing, which is exactly why people end up with 30 hours of audio they never open. A recording you process into 9 cards is worth more than 10 recordings you don't. Ask your tutor before recording either way.


Sources

Why spacing is the part that matters

Why building the card yourself helps

Why the audio matters