Verbamor Download the appGet the app
Free tools Download on the App Store
Free tool

How many Spanish words do you know?

Tick the words you'd recognise in a sentence. 25 of the 97 are fake, and how often you tick those is what stops this tool flattering you. The estimate is band-weighted, and the arithmetic is printed underneath so you can check it.

Most vocabulary tests on the internet ask you to tick words you know and then multiply. That gets you a number, and the number is too big.

Two things are wrong with it. Ticking is cheap, so people tick words they half-recognise, and nothing in the test catches that. And a flat percentage treats every word as equally hard, when the words in a frequency list get very much rarer as you go down.

This one fixes both. It draws 12 words from each of 6 frequency bands, weights each band by how many ranks it covers, and mixes in 25 pseudo-words, 4 or 5 to a band, to measure how often you tick something that isn't there.

The test

Tick every word you'd understand if you met it in a sentence. Don't look anything up, and don't linger. Some of these words are invented, and ticking them is how the tool learns to discount you.

Every word, every fake and every rank is in the page source. JavaScript only does the adding up, so with it switched off you still get the full test, the formulas and a worked example, which is enough to score this by hand. The complete list with ranks is in the sources section at the bottom.

Ranks 1 to 500 Some of these are invented. Tick only the ones you would understand in a sentence.
Ranks 501 to 1,500 Some of these are invented. Tick only the ones you would understand in a sentence.
Ranks 1,501 to 3,000 Some of these are invented. Tick only the ones you would understand in a sentence.
Ranks 3,001 to 6,000 Some of these are invented. Tick only the ones you would understand in a sentence.
Ranks 6,001 to 12,000 Some of these are invented. Tick only the ones you would understand in a sentence.
Ranks 12,001 to 25,000 Some of these are invented. Tick only the ones you would understand in a sentence.

Your estimate, with the working

Tick some words and press the score button. Three figures appear below, and the table under them shows every multiplication that produced the total.

How the number is built

Three steps, and the middle one is the whole point.

Step 1, the hit rate per band. Each band has 12 real words in it. Ticking 9 of them gives a hit rate of 0.75 for that band, and nothing else in the band matters yet.

Step 2, subtract your guessing. Your false-alarm rate f is the share of the 25 pseudo-words you ticked. Brysbaert and colleagues, working with 221,268 people in 2016, state the correction as plain subtraction: they took the proportion of yes-responses to the non-words away from the proportion of yes-responses to the words. This tool does that and then rescales, which is the standard high-threshold form:

corrected = (hit rate − f) ÷ (1 − f)
f = pseudo-words ticked ÷ 25
any corrected rate below 0 is clamped to 0

The division is why a perfect scorer still reads 1.0. Plain subtraction pushes even someone who ticks every real word down to 1 − f, which under-reads at the top of the range. Both figures appear in the panel above, so you can see the size of the gap on your own answers.

Step 3, weight each band by its own width. The bands are not the same size. The first covers 500 ranks, the last covers 13,000. So a band's corrected rate is multiplied by the number of ranks it stands for, and the six products are added:

estimate = Σ ( corrected rate × band width )

band 1–500 → width 500
band 501–1,500 → width 1,000
band 1,501–3,000 → width 1,500
band 3,001–6,000 → width 3,000
band 6,001–12,000 → width 6,000
band 12,001–25,000 → width 13,000

A worked case, so you can check the code against a pencil. Say you tick 12, 11, 9, 6, 3 and 1 across the six bands, and 5 of the 25 pseudo-words, so f = 0.20.

band 1: (1.0000 − 0.20) / 0.80 = 1.0000 × 500 = 500.00
band 2: (0.9167 − 0.20) / 0.80 = 0.8958 × 1000 = 895.83
band 3: (0.7500 − 0.20) / 0.80 = 0.6875 × 1500 = 1031.25
band 4: (0.5000 − 0.20) / 0.80 = 0.3750 × 3000 = 1125.00
band 5: (0.2500 − 0.20) / 0.80 = 0.0625 × 6000 = 375.00
band 6: (0.0833 − 0.20) / 0.80 = negative, clamped to 0 × 13000 = 0.00

total = 3,927 words

The same ticks with no correction come to 6,625, so skipping the pseudo-words would have added 2,698 words, or 69%. And the naive version, 42 ticks out of 72 times 25,000 ranks, gives 14,583. That last figure is nearly 4 times the honest one, and it is what most of these tests print.

The sibling post on how many words a native speaker knows is where this correction comes from, and it also shows what happens when the field skips it: Milton and Treffers-Daller re-ran Goulden's own test items with every tick marked for correctness, and the answer fell 42.9%.

What the number does not mean

Being blunt about the edges, because a vocabulary figure with no method attached is worth nothing.

The unit is surface forms, not lemmas or word families. The source list counts hablar and hablo as separate ranks. That makes the number look larger than a lemma count of the same knowledge, and much larger than a word-family count. Comparing it against a figure from a different unit is a mistake, and it's the mistake behind most of the disagreement in the published literature.

Everything past rank 25,000 is invisible. The test samples nothing out there, so the estimate can never exceed 25,000 no matter how good your Spanish is. If you tick all 72 real words and none of the fakes, you get 25,000 and the tool has stopped measuring you.

72 items is a small sample. A single band is 12 words, so one extra tick moves that band by 8.3 percentage points, which in the last band is 1,083 words of estimate. Treat the result as a rough band, not a figure to quote. For reference, Goulden's much-cited 17,200 came from 20 people at 250 items each, and Brysbaert's came from about 14.8 million judgements.

Ticking is recognition, and shallow recognition at that. Brysbaert says of his own equivalent numbers that they "should be considered as upper estimates of word knowledge, going for width rather than depth". You can tick a word whose second sense you've never met. Productive knowledge, the words you can actually say, is thought to run at roughly half of receptive.

The corpus is film and TV subtitles. Frequency ranks from subtitles skew toward speech, which suits a conversational learner and misreads an academic one. A list built from newspapers would rank these words differently. Every published Spanish, French, Italian and Portuguese frequency list compares the options, including which ones you may legally use.

The words were picked by hand out of the bands. A random draw would be cleaner. I filtered proper nouns, single letters and obvious loanwords by eye, which is a judgement call a mechanical sample would not need. The full list with every rank is below, so you can check each one against the source file yourself. Score the test first and every word shows its rank, with the invented ones named.

Sources and licence

The word list is the 2018 Spanish file from hermitdave/FrequencyWords, built from OpenSubtitles via OPUS. The repository's README licenses the code MIT and the content CC BY-SA 4.0, and the 97 words below are a derivative of that content under the same licence. Each rank below is that word's line number in es_50k.txt, checked on 4 September 2026. Here is the whole sample:

Every word in the test, by rank band, with its rank in the source frequency list, plus the 25 invented words.
BandWords, with rank
1 to 500señor 107, hombre 121, noche 136, mañana 173, nombre 216, cuenta 262, historia 314, semana 326, punto 393, cuerpo 435, número 458, lista 490
501 to 1,500fuego 603, cuarto 628, ropa 652, noticias 675, cambiar 699, gusto 755, personal 778, piso 1,007, juicio 1,129, hogar 1,154, esperanza 1,252, vuelo 1,303
1,501 to 3,000proyecto 1,573, príncipe 1,610, serie 1,649, correo 1,875, cerrado 1,914, techo 2,103, reino 2,327, campaña 2,366, taza 2,443, despacho 2,520, análisis 2,628, patio 2,963
3,001 to 6,000ambiente 3,001, piano 3,149, pensamiento 3,224, lenguaje 3,598, emoción 3,745, pasaporte 4,119, fianza 4,344, lata 4,420, condena 4,572, remedio 4,716, oreja 4,942, tristeza 5,094
6,001 to 12,000sillas 6,308, vestuario 6,611, cerradura 6,779, fotógrafo 6,906, comisionado 7,208, contenedor 7,663, gabinete 8,554, tubos 9,457, inventario 9,755, estrecho 10,951, pizca 11,107, impulsos 11,400
12,001 to 25,000agricultura 12,001, natación 12,636, paternidad 12,959, bistec 13,279, encaje 13,605, cordura 14,566, varones 15,209, noticiero 15,866, gurú 16,820, debut 17,148, colegios 17,466, titanio 22,045
Inventedcurbeza, gralofa, plandar, rusteño, talvento, dorreta, misgano, polantre, zurbeda, hentira, brasculo, nadrino, carpeño, luvante, tresmillo, ordicia, fambusa, gomerato, trabundo, melquiso, pardiseo, sunterna, valpiche, cerdumbo, nostalvo

The correction comes from Brysbaert, M., Stevens, M., Mandera, P., & Keuleers, E. (2016). How Many Words Do We Know? Frontiers in Psychology, 7:1116. Their test ran 67 words against 33 non-words per session, and their subtraction is quoted in full in the post on native-speaker vocabulary size.

The 25 pseudo-words are mine. Each follows Spanish phonotactics and none appears anywhere in the 1,202,520-type full OpenSubtitles Spanish list, which is the strongest check I can run without a dictionary API. That is not proof no Spanish speaker has ever used one of them.

The rest of the frequency argument, and why the first 650 words do so much of the work, is in the word-frequency curve.

Counting words is the easy part.

Verbamor counts the ones you actually stumbled on in your own lessons, builds the cards with native audio, and brings each one back when your recall has sagged to 90%.

Download on the App Store