Recording your first lesson so the transcript is usable
A phone lying face down on a café table, six feet from your tutor, records a lesson that comes back as mush. The words you needed are in there; the transcription engine cannot hear them over the espresso grinder. A lesson happens once, so a bad capture costs you the whole hour.
A phone lying face down on a café table, six feet from your tutor, records a lesson that transcribes into mush. The words you needed are in there. The machine cannot hear them over the espresso grinder, and what comes back is a wall of half-guessed text with nothing worth carding in it.
A lesson is a live event. You cannot go back and record it again, so a bad capture costs you the whole hour and everything you said in it.
Put the phone in the right place
Verbamor captures at 16 kHz mono, which is the setting speech runs at. Transcription models and the speaker-separation models behind them both operate at 16 kHz, so a higher sample rate gets discarded on the way in. Microphone placement is therefore the only audio decision you get to make.
Three placements, in order of how well they work.
- In-person lesson. Phone face up, screen up, on the table between you and your tutor, roughly an arm's length from each of you. Not in a pocket, not in a bag, not face down. The iPhone microphones sit at the bottom edge and the top edge, and a table surface pressed against them removes most of what you want.
- Video call on a laptop. Put the phone beside the laptop speaker, face up, about a hand's width away. You will capture your tutor through the speaker and yourself through the air. Both arrive.
- Video call on the same phone. This one does not work. iOS gives the microphone to one app, and a call app holds it. Take the call on a laptop or a second device, and record with the phone.
I have not measured the exact distance at which speaker labels start to break, so I will not give you a number I did not test. The published range is a useful anchor. The AliMeeting corpus, 120 hours of recorded meetings built for the 2022 ICASSP transcription challenge, states that its microphone-to-speaker distance runs from 0.3 to 5.0 metres, roughly one foot to sixteen feet. An arm's length sits at the easy end of that range and a café table away sits at the hard end. Use the one-minute test below to judge your own room.
Kill the obvious noise sources first. Human ears filter out a running tap, a fan pointed at the table, or background music, and transcription models do not, which is why a recording that sounded fine to you comes back wrong.
Say the sentence to your tutor before you press record
You are recording another person. Tell them, in advance, in one sentence. This is the wording I use:
"I record our lessons so I can make flashcards from them afterward. Is that alright with you?"
Most tutors say yes immediately, and a good number will start speaking more clearly once they know. If one asks what happens to the audio, the honest answer is that the file is uploaded, transcribed, and stored on your account.
Two situations where you should not record. A tutor who says no. And any jurisdiction where recording a conversation needs consent from everyone in it, which is the law in several US states and much of Europe. Asking first solves both problems. If the answer is no, paste your notes in as text and write by hand during the lesson.
The four limits that will bite you
iPhone, iOS 16.4 or later. The app is iPhone-only right now, so an iPad will run it at phone size and a laptop will not run it at all. The microphone permission prompt appears the first time you record, and if you dismiss it, iOS will not ask again: turn it back on in Settings, using the Open Settings button the app puts on the error.
Ninety minutes per recording. This is Verbamor's own cap, not a restriction imposed by anything upstream. The transcription runs on ElevenLabs Scribe, whose documented ceilings are 10 hours and 3 GB per file, far above anything a lesson produces. Verbamor refuses past 90 minutes because Scribe bills per second of audio, and a recorder left running in a pocket is the one way to run up an unbounded charge. A typical 50-minute lesson fits with room left over, and a two-hour class needs two recordings. A second Verbamor limit sits behind it: 450 minutes of transcription per account per day.
Three minutes of silence ends it. If the app hears nothing above roughly minus 40 dBFS, which is below the level of someone talking quietly across a table, for three continuous minutes, it stops the recording and keeps what it has. This exists because the common failure is a lesson that ends, a phone that goes in a pocket, and a recording that keeps running for an hour. Ordinary thinking time and looking things up will not trip it.
Fifty megabytes per uploaded file. This one is not Verbamor's choice either: 50 MB is the fixed upload ceiling on the Supabase plan the backend runs on, and the app checks against it before uploading rather than after. Verbamor's own captures never come close. At 16 kHz mono and 32 kbps, an hour of its audio is about 14 MB, and a full 90-minute recording lands around 21 MB. Imports are where you will meet the limit, because a file recorded elsewhere at 48 kHz stereo is roughly 20 times denser per minute. A 30-minute music-bitrate m4a can exceed 50 MB on its own.
The screen can lock while you record. Background capture is supported on purpose, so start the recording and let the screen go dark rather than burning battery for a full lesson.
When to skip the microphone entirely
If your lesson happens on a platform that records for you, import that file instead. It is almost always cleaner than a phone on a table, because each speaker was captured at their own microphone. Import accepts any audio file already on the phone, including a recording your tutor sent you after the call.
When no audio exists at all, paste text instead. A chat transcript or a page of lesson notes runs through the same vocabulary mining with no audio involved. Paste needs about a dozen words minimum, so a two-word note comes back empty.
The one-minute test after the lesson
- Open the recording. Read seven lines of the transcript from the middle of the lesson, not the start.
- Count the label errors. A label error is a line where the speaker label did not change even though the voice did, or changed when it did not. The clearest tell is a correction sitting under the same label as the mistake it corrects, because that means one person's error and another person's fix have been merged into one turn. Two or more label errors in seven lines means throw the recording away and move the phone closer next time.
- Look for your own worst moment. Find a place where you struggled to say something. If it is legible, the capture is good.
- If the labels held but individual words are misspelled, keep the recording. Mine it, then skip any item whose spelling you cannot confirm against a dictionary or the tutor's own correction. A misspelled word on a card teaches you the misspelling, and the labels are what tell you who said which version.
The two transcripts below are not one lesson captured twice, which one phone cannot do. The first is a real transcript from a 48-minute Spanish lesson recorded in June 2026, phone face up on the table, an arm's length from each speaker:
Teacher: ¿Y qué hiciste el fin de semana?
Student: Fui a la playa con mi hermana. Estaba... eh... estaba muy lleno.
Teacher: Había mucha gente, sí. Se dice "había mucha gente".
Student: Había mucha gente. Vale.
The labels alternate on the turns, the Spanish is spelled correctly, and the one stumble that matters is preserved: the learner reached for estaba muy lleno and the tutor supplied había mucha gente. That exchange is a card.
The second is what that same audio came back as after we played it in a café and re-recorded it from six feet away, phone face down, next to the espresso grinder:
Speaker 1: ¿Y qué hiciste el finde se va?
Speaker 1: Fui a la playa con mi ermana esta ba eh
Speaker 2: hay mucha gente si se dice
Speaker 1: ay mucha gente bale
Run step 2 on it. Lines one and two are both Speaker 1 when the voice changed, and line four puts the learner's attempt and the tutor's correction under one label, which is the merged-turn tell. That is two label errors in four lines, so it fails before you read the spelling. The spelling confirms it: hermana lost its h, vale became bale, and había collapsed to hay and then ay, so the grammar point the lesson was about is gone.
Then pick your items. Ten to fifteen from a first lesson is plenty, and the app can surface up to 60 words and 30 phrases from one recording. Taking all ninety is how you end up with a study queue you avoid opening. Multi-word phrases come back whole, so tener ganas de arrives as one item rather than three unrelated words, and phrases are usually the better keep because they carry the grammar with them.
The best single filter: keep the word if you tried to say it during the lesson and could not. The recording is the only record of where that gap was, and the failed attempt is itself worth something. Kornell, Hays and Bjork tested that in 2009 across six experiments using materials designed so the guess would fail, and found that a wrong attempt followed by the answer beat being handed the answer straight. The word you reached for and missed is already half learned. Words your tutor used that you understood fine are much weaker keeps, because understanding one is not the same skill as producing it.
Where this advice does not apply
If your lessons are grammar drills from a textbook, recording them is close to useless. The transcript will be full of conjugation tables read aloud, and the vocabulary worth keeping is already printed in front of you. Type those in through Paste text and skip the microphone entirely.
The same goes for a lesson conducted mostly in English. The mining looks for the language you are learning, and forty minutes of an English explanation of the subjunctive produces very little to card.
And if you have no recurring source of real language at all, no tutor, no class, no conversation partner, then the recording setup is the wrong problem to be solving. Fix the input first. The app turns language you already meet into language you can produce, and with nothing coming in there is nothing for it to work on.
Before your next lesson: set the phone face up on the table, ask your tutor the one-sentence question, and start the recording before the small talk ends. The small talk is where the useful words hide.
Where you put the phone decides everything.
Verbamor transcribes and mines whatever you record, but only if the audio holds up, so an arm's length from your tutor is the whole difference between a working deck and a wall of guessed text.
Start 7 days free