Verbamor Download the appGet the app
Free tools Download on the App Store
Free tool

How much of this text is already inside your first 5,000 words?

Paste something in Spanish, French, Italian or Portuguese. You get the share of its running words that sit inside the first 1,000, 2,000, 3,000 and 5,000 frequency ranks, the list of words that fall outside, and a proper-noun count kept out of the headline figure on purpose.

Coverage is the share of running words in a text that you already know. It's the number behind every "learn 1,000 words and you can read a newspaper" claim, and it's worth measuring on a text you actually want to read rather than on somebody's corpus.

So paste one in. Everything runs in your browser and nothing is sent anywhere.

The checker

It counts forms, and that matters

This tool counts written forms. Hablo and hablas are 2 separate items in the list and 2 separate lookups here.

The famous coverage targets are not counted that way. Paul Nation's 2006 paper, the source of "8,000 words to read a novel," counts word families: a headword plus every inflection and derivation you could decode from it. His own lists average 4.33 forms per family, so 8,000 families works out at 34,660 distinct written forms.

Which means a percentage from this page and a percentage from that paper are different measurements wearing the same sign. If you hit 90% here at 3,000 forms, you have not hit Nation's 3,000-family mark. Romance verbs make the gap worse: an English verb has 4 inflected forms, and hablar has around 44 before a derivation like hablador joins it.

Lemmatising would fix the unit, and I didn't do it. A tagger that guesses wrong on porto, which is a Portuguese city and a form of portare, corrupts 2 ranks at once, and I'd rather show you a number whose flaw is printed than a number whose flaw is hidden inside a model. So: forms.

Why names get their own row

Proper nouns are coverage you get for free, because Madrid is Madrid in every language you'll read it in. They are also additive to the published figures rather than part of them. Nation's conclusion says the first 1,000 words plus proper nouns cover 78% to 81% of written text, and the clause falls off almost every time somebody quotes the number.

The block is not small and it is not stable. Across the 14 texts Nation profiled it runs from 0.50% of The Turn of the Screw, a ghost story with 4 characters in it, to 6.12% of the Brown newspaper corpus. His 5 novels average 1.45% and his 5 newspaper corpora 5.41%, and the 2 groups don't overlap at all.

So a checker that quietly folded names into one percentage would move your score by up to 5 points depending on whether you pasted a novel or a news story. This one counts them, prints them, and leaves them out of the band figures.

The list, and its licence

The ranks come from hermitdave/FrequencyWords, built from the OpenSubtitles 2018 corpus through OPUS. The top 5,000 forms per language, lowercased, with duplicates removed after lowercasing. Downloaded 4 September 2026.

The licence is CC BY-SA 4.0 for the data. The repository's README splits that from the MIT licence covering its code, and the LICENSE file holds only the MIT text, which is how "hermitdave is MIT" got started. Share-alike is a real obligation and this page honours it by attribution.

It's the list I could ship. Of the free options covering all 4 of these languages, it's the only one whose licence permits commercial use. SUBTLEX-ESP and SUBTLEX-IT are CC BY-NC-SA, SUBTLEX-PT-BR adds NoDerivatives on top, SUBTLEX-PT carries no licence file at all, and KELLY Italian is CC BY-NC-SA 2.0. I went through every list and read each licence at its own source, and several of the most-recommended ones would have been the wrong answer here.

Being subtitles, the corpus is dialogue. These ranks describe how people speak on screen, so a legal contract or a chemistry paper will score badly and that is the corpus talking, not you.

What this will get wrong

The proper-noun test is a capital letter that isn't at the start of a sentence. It's a heuristic, and it has 5 failures worth knowing. A name in the first position of a sentence gets counted as vocabulary. A text pasted in all lower case gets no names at all, and one pasted in ALL CAPS gets almost every word binned as a name. A capitalised common noun mid-sentence gets wrongly binned. And an abbreviation ends a sentence as far as this is concerned, so in Sr. Garcia fue a casa the period after Sr makes Garcia look sentence-initial and it slips through as vocabulary. The count is printed so you can eyeball whether it looks right for what you pasted.

Anything past rank 5,000 shows as outside. The list itself runs to 50,000, and I cut it at 5,000 because that's the range a learner works in and 4 full lists would be a 12 MB page.

Numbers, and tokens with no letters in them, are skipped entirely rather than counted as unknown. Apostrophes and hyphens follow the list rather than a general rule, because the list treats them differently. An elided clitic keeps its mark and stands alone, so l'homme is l' plus homme, and l' is rank 16 in French. A hyphenated compound stays whole, so est-ce is one word at rank 70 and peut-ĂȘtre one at 130. Portuguese carries 206 hyphenated forms like diz-me on the same rule. Splitting both marks, which is what this did first, put 202 French and 215 Portuguese entries permanently out of reach and sent aujourd'hui to the unknown pile from rank 271.

And the whole thing measures recognition of shapes on a page, which is the easiest thing vocabulary does. Reading desarrollo and knowing it means development is not the same skill as producing it in a sentence while somebody waits. The frequency curve explains why the first thousand ranks buy so much and the tenth buys so little.

Every load-bearing claim Verbamor makes is traced to its paper on the research page.

Knowing a word is the easy half.

Verbamor takes the words you got stuck on in your own lessons and schedules them back with native audio, so recognition turns into recall.

Download on the App Store