One research team tested 10 rival language apps
Duolingo, Babbel, Busuu, Rosetta Stone, Pimsleur, Mango, italki, LingQ, Language Zen and Hello English all point at efficacy research by the same statistician. Sixteen reports since 2009. I read all 16, and the surprise is not what you would guess.
Open Duolingo's old marketing and you get "an independent study found that Duolingo trumps university-level language learning." Open Babbel's and you get an efficacy study. Busuu, the same. Rosetta Stone, Pimsleur, Mango, italki, LingQ.
Follow each one to its source and you arrive at the same place: Roumen Vesselinov and John Grego.
The reports live at comparelanguageapps.com, which credits its design to Vesselinov and lists him as sole contact. I downloaded all 16 PDFs on 3 September 2026 and read them. I went in expecting a funding problem, and what I found moves the criticism elsewhere.
Sixteen reports, 10 clients, one team
The site's research index, author lines as printed. Vesselinov is on all 16, Grego on 15.
A note on the count, because I have seen 14 quoted and could not get there. The index lists 16. Distinct app companies that commissioned work: 10. Rosetta Stone, Duolingo, Language Zen, Babbel, Busuu, Hello English, italki, Pimsleur, Mango, LingQ. Add Emmersion Learning, whose TrueNorth test the team validated in 2020, for 11 clients. Auralog and Berlitz appear as compared products in 2009 without commissioning anything, so they stay out. The homepage claims "20+ Years of Research" since 2005; the oldest report I can open is January 2009.
One research team, then, has produced the efficacy evidence for most of the language app market over 16 years, and the apps compete.
They do disclose the funding
Here I correct my own starting assumption. I expected reports that stayed quiet about who paid. Instead 15 of 16 name their funder, usually page 5 or 6, under study limitations, in wording that barely changes across a decade. Duolingo's 2012 report:
This study was funded by Duolingo but the data collection and the analysis were done independently by the Research team.
Vesselinov & Grego 2012, Duolingo Effectiveness Study, page 5Babbel 2016: "This study was funded by Babbel." Busuu 2016: "funded by busuu." italki 2018, Mango 2019, Pimsleur 2019 and Rosetta Stone 2019 run the same sentence with the name swapped. LingQ 2023 and Busuu 2025 shift to "The cost for this study was covered by" the client. Hello English 2017 breaks the pattern in a good way: its funder was the Central Square Foundation, a grant-making organisation, not the app.
The 2021 Busuu report goes furthest, filing the funding under "Limitations of the Study" and naming a second problem itself:
This study was funded by Busuu and participants were existing users of Busuu who were incentivized to take part in the study through the provision of free language lessons. Both of these limitations may have led to a degree of bias.
Vesselinov, Grego, Tasseva-Kurktchieva & Sedaghatgoftar 2021, page 26That is a better conflict statement than plenty of published journal work carries. The rest of this post is critical, and criticism counts only if the credit is real.
The reports also name the structural problem before I get to it. Babbel 2016 asks for help "to require the creators of language learning apps to provide independent efficacy measures." Busuu 2021 is blunter: there are "currently very few opportunities for language app developers to benefit from independent research studies into their efficacy, and so the current best way for them to give confidence to their customers is through commissioning research." The commissioned researchers are saying it is the only game available.
The exception is the oldest. The January 2009 Rosetta Stone report, out of Queens College CUNY, carries no funding sentence: I searched the text for fund, sponsor, commission, support and grant and got zero hits. Its two 2009 companions on motivation both say the work "was commissioned by Rosetta Stone." One out of 16, and the oldest, is the honest way to say it.
Krashen made the Duolingo funding point in January 2014, in the International Journal of Foreign Language Teaching: the study Duolingo cited "is Vesselinov and Grego (2012), funded by Duolingo." He was reading it off the report.
The league table drops the column
The actual problem sits one layer up. The hub also publishes two ranking tables, vocabulary and grammar plus oral proficiency, putting the apps in a single list against each other. The vocabulary table ranks eight apps by "Study Hours for First College Semester," lowest first.
Every row comes from a study paid for by the company in that row. The columns are year, app, percent improved, confidence interval, mean hours, range, sample size and a report link. No column for who paid, and the words fund and sponsor appear nowhere on either page.
Nothing is hidden: click through and the funder is on page 5. But the table is the artifact that gets screenshotted, and a reader comparing 13 against 34 has no signal that these are 8 separate commissions.
And the site invites you to read it as a bake-off. The 2025 Busuu report says the statistical design "and methodology are comparable for all studies," and LingQ 2023 says the same. That claim turns 16 client reports into a ranking, and does the heaviest lifting on the site.
I am not the first to doubt that, and the doubt comes from an unlikely direction. A 2021 paper in Foreign Language Annals (doi:10.1111/flan.12600), four of whose five authors list Duolingo as their affiliation, notes the reports "were published on company websites as white papers, not peer-reviewed journal articles," and that because pretest scores differ, "it is hard to compare the effectiveness across these products." A peer-reviewed journal says the league table cannot carry the weight, and the people saying so work for a company sitting in it.
Duolingo sits at the bottom with 34 hours, on the oldest study in the set, run in 2012 on a product version that no longer exists. LingQ sits at the top with 13, from 2023. Whether a 2012 app and a 2023 app belong in one league table is the question the table does not ask.
Who finished varies by 50 points
Here is the arithmetic that convinced me. Each report states its initial random sample and its final study sample. Divide one by the other for a completion rate. Neither the reports nor the tables put those side by side, so I did.
The spread is 50.1 points, 43.7% to 93.8%. A Language Zen participant was 9 times likelier to drop out than a Hello English one. Strip out Hello English, which recruited school students rather than app users, and the other 14 still spread 40.7 points.
Pooled across all 15, 4,330 entered and 3,129 finished: 28% of starters absent.
That matters for the ranking. LingQ tops the table at 13 hours, having lost 47.4% of its starters. Duolingo sits last at 34, having lost 55.1%. Babbel is mid-table at 21 and kept 83.1%. The table shows sample size, never what the sample started as.
The oral table has a sharper version. Busuu's 2016 oral row reports 75% improving a level, on 61 people, a subsample of the 144 selected for having done "about 16 hours of study or more." That describes the people who studied most: defensible as analysis, poor as a thing to rank other apps against without a note.
The reports are careful here. Most compare completers against dropouts on age, gender, education and initial test score, and report no significant difference. That checks demographics. It says nothing about the variable deciding the result: whether you keep using the app.
Krashen caught the same shape from the other end. Of the 156 he counts as starting, 90 lasted to the end, 88 had usable scores, and 66 filled in the exit survey where 78.8% said they were satisfied. That is 66 people out of 156.
One more from the Duolingo report, the most quoted of the set. The famous 34 hours comes from a mean gain of 8.1 WebCAPE points per hour. The same report prints the median: 3.9. Redo the division on the median and 270 points, the second-semester cutoff, needs about 69 hours. Same study, same page, and the answer doubles depending which average you pick.
And every hours figure in that table is a mean. Duolingo is the only study printing both, where the mean flatters by 2.1 times. If that is near typical, the whole column is optimistic by an unknown factor.
Which way does the attrition cut? These studies keep only people who did at least 2 hours and completed both tests, so survivors are the persistent users, and heavy attrition should make an app look more efficient per hour. That predicts high-dropout studies sit near the top. LingQ lost 47.4% and tops the table at 13 hours, which fits. Duolingo lost 55.1% and sits last at 34, which does not. Two points facing opposite ways is not a finding, and across 15 studies with differing tests, years and recruitment I cannot separate selection from the rest. Neither can the table.
What to do with a study like this
This involves named people publishing under their own names, so I will be exact about which is which.
What is documented: one team produced 16 efficacy reports for 10 competing app companies between 2009 and 2025. Fifteen name their funder; the oldest carries none I could find. The hub publishes two tables ranking those apps against each other, neither with a funding column. Completion rates range from 43.7% to 93.8%.
What is inference, and mine: a table built from 10 separate commissions, spanning 13 years of product versions and a 41-point spread in who finished, is weaker than it looks. The funder belongs next to each row, the completion rate beside the sample size.
What I am not saying: that any result was bought, altered or faked. I have no evidence of that and went looking for none. Funded research is normal here and everywhere else. These reports state their funding and describe their methods in enough detail that a stranger can check the arithmetic, which I did. More than most marketing claims allow.
So, three questions for anyone reading "scientifically proven" on a pricing page. Who paid, answerable in 90 seconds by searching the PDF for funded. How many started, against how many finished. And how old the tested product is, because a 2012 app and the app on your phone share a name and little else.
Worth knowing what gets measured, too. Almost all use WebCAPE, a multiple-choice Spanish placement test, some with an oral interview alongside. That is vocabulary and grammar recognition, some distance from conversation. The strongest published result in this category is still a receptive one.
Verbamor has no efficacy study. Commissioning one costs more than the business makes, so read that as a disclosure. What it does is narrower than anything in that table: it turns words from lessons you already had into retrieval practice scheduled by FSRS. Those are claims about spacing and retrieval, with a research base that has nothing to do with me. I have never measured Verbamor against Babbel, and neither has anyone else. For the per-hour comparison, see that one with the numbers shown.
Sources
The hub and its tables
- Compare Language Apps. comparelanguageapps.com, research index and homepage. Lists 16 reports, 2009-2025, and states "16+ Language Apps Tested" and "20+ Years of Research" since 2005. Retrieved 3 September 2026.
- Compare Language Apps. Vocabulary/Grammar Proficiency Results. Beginner rows: LingQ 2023 (13 hours, n=101), Rosetta Stone 2019 (13, n=143), Mango 2019 (15, n=87), italki 2018 (19, n=102), Babbel 2016 (21, n=325), Busuu 2016 (22, n=144), Language Zen 2015 (25, n=101), Duolingo 2012 (34, n=88). Table headers extracted from the page HTML: Year, Language App, Improved Proficiency (Percent, 95% CI), Study Hours for First College Semester (Mean, Range), Sample Size, Statistical Report. No funding column, and a case-insensitive search of the page source for "fund" and "sponsor" returns nothing. Retrieved 3 September 2026.
- Compare Language Apps. Oral Proficiency Results. Headers: Year, Language App, Improved Proficiency One Level Up or more (Percent, 95% CI), Improved Proficiency Two Levels Up or more, Sample Size, Statistical Report. Same absence of any funding column or the words fund and sponsor. Retrieved 3 September 2026.
The funding statements, quoted from the reports
- Vesselinov, R. & Grego, J. (2012). Duolingo Effectiveness Study. Page 5: "This study was funded by Duolingo but the data collection and the analysis were done independently by the Research team." Initial random sample 196, final study sample 88; mean gain 8.1 WebCAPE points per hour, median 3.9.
- Vesselinov, R., Grego, J., Tasseva-Kurktchieva, M. & Sedaghatgoftar, N. (2021). The Busuu Efficacy Study 2021. Under "Limitations of the Study": "This study was funded by Busuu and participants were existing users of Busuu who were incentivized to take part in the study through the provision of free language lessons. Both of these limitations may have led to a degree of bias." Initial sample 141, final 119, stated dropout 15.6%.
- Same sentence pattern verified in: Babbel 2016 ("funded by Babbel"), Busuu 2016 ("funded by busuu"), Language Zen 2015, italki 2018, Mango Languages 2019, Pimsleur 2019, Rosetta Stone 2019; LingQ 2023 and Busuu 2025 use "The cost for this study was covered by"; TrueNorth 2020 states Emmersion Learning "provided the data and funding for this study"; Hello English 2017 names the Central Square Foundation. All 16 PDFs downloaded from comparelanguageapps.com and text-searched on 3 September 2026.
- Vesselinov, R. & Grego, J. (2016). The Babbel Efficacy Study. "More help is needed from users, investors, and analysts to require the creators of language learning apps to provide independent efficacy measures." The same sentence appears in the Mango Languages 2019 and Pimsleur 2019 reports. Busuu 2021 states there are "currently very few opportunities for language app developers to benefit from independent research studies into their efficacy, and so the current best way for them to give confidence to their customers is through commissioning research with engaged groups of users."
- Vesselinov, R. (2009). Measuring the Effectiveness of Rosetta Stone. Queens College, City University of New York, January 2009. A full-text search for fund, sponsor, commission, support and grant returned no funding statement. Its two 2009 companion reports both state the work "was commissioned by Rosetta Stone."
The comparability claim
- Vesselinov, R., Grego, J. & Tasseva-Kurktchieva, M. (2025). Busuu 2025 Six-Language Efficacy Study. "Our previous studies evaluated Rosetta Stone, Duolingo, Busuu, Babbel, Mango Languages, Pimsleur, Hello English, italki, and Language Zen. Statistical design and methodology are comparable for all studies." Initial sample 1,635, final study sample 1,205. The report states "430 people (28.1%) did not complete," but 430/1,635 = 26.3%; the headcount is consistent with the samples and the printed percentage is not.
- Vesselinov, R., Grego, J. & Tasseva-Kurktchieva, M. (2023). LingQ Efficacy Study 2023. Same comparability sentence. Initial random sample 192, final 101, stated dropout 47.4%.
The completion arithmetic
- Vesselinov, R. & Grego, J. (2016). The busuu Efficacy Study. "The final study sample for written proficiency consisted of 144 people... The final subsample for oral proficiency was part of the 144 people sample and consisted of 61 people with about 16 hours of study or more and valid initial and final OPIc tests. The mean study time for the oral test sample was about 24 hours." The Oral Proficiency table publishes the 61 as its sample size without the 16-hour selection rule.
- Initial random sample and final study sample, both as printed in each report: Rosetta Stone 2009 effectiveness 176/135, Rosetta Stone 2009 motivation 203/164, Rosetta Stone/Auralog/Berlitz 2009 comparative 303/234, Duolingo 2012 196/88, Language Zen 2015 231/101, Babbel 2016 391/325, Busuu 2016 196/144, Hello English 2017 97/91, italki 2018 125/102, Mango 2019 149/95, Rosetta Stone 2019 175/143, Pimsleur 2019 120/82, Busuu 2021 141/119, LingQ 2023 192/101, Busuu 2025 1,635/1,205. Completion rates 76.7, 80.8, 77.2, 44.9, 43.7, 83.1, 73.5, 93.8, 81.6, 63.8, 81.7, 68.3, 84.4, 52.6 and 73.7 percent. Spread 50.1 points, from Language Zen’s 43.7 to Hello English’s 93.8; excluding Hello English, which recruited school students rather than app users, 40.7 points across the other 14. Pooled 3,129 of 4,330 = 72.3%, so 27.7% absent. Every one is my division of their published figures; the site publishes no completion comparison. The 2020 TrueNorth report is a psychometric validation of a test and publishes no comparable initial/final pair, so it is the one report of 16 absent from this column. Note Krashen counts 156 starters for the Duolingo study, the sample before 8 too-advanced participants and 7 refusals were removed; using the report’s own post-exclusion figure of 196 keeps the basis consistent and is the harsher number for Duolingo.
The outside critique
- Krashen, S. (2014). Does Duolingo "Trump" University-Level Language Learning? International Journal of Foreign Language Teaching, January 2014, pages 13-15. States the study "is Vesselinov and Grego (2012), funded by Duolingo"; reports 90 of 156 starters lasting to the end, mean time 22 hours with SD 20.4, and rate by reason for study: travel 17.6 (n=10), business/work 11.4 (n=16), personal interest or school 5.7 (n=62). Acknowledges "very helpful and patient discussion" with Vesselinov.
Every load-bearing claim Verbamor makes is traced to its paper on the research page.
No efficacy study. Just the words from your own lessons.
Verbamor turns what your tutor taught you into scheduled retrievals, so the vocabulary survives the week.
Download on the App Store