Prior academic work checking the prevalence of hallucinations in pre-print servers tends to estimate that on the order of 1% of citations are AI hallucinations — and growing. Zhao et al. (2026) put it at 0.39% on arXiv, 0.21% on bioRxiv, 1.91% on SSRN, and 0.27% in PubMed Central as of August 2025. Topaz et al. (2026) in The Lancet found a twelve-fold rise in fabricated references in biomedical papers since 2023. I think these are likely to be an underestimate if anything, based on my work at VerusCite.
The reason is that the large scale audits (see the several examples cited in my benchmark) only look for differences in titles in online academic article databases (like OpenAlex).
That design has three consequences:
- A real title with completely wrong authors is a clean match
- A real paper placed in a journal it never appeared in is a clean match
- News articles, agency reports, statutes, working papers, books, and bare URLs are dropped from the analysis rather than checked
What VerusCite does differently
My tool has several differences from these large scale analyses. These include:
- It compares authors and publication name, not just the title. A real paper credited to the wrong people, or placed in the wrong journal, gets flagged.
- It checks every reference type. Books, chapters, conference papers, working papers, court cases, news articles, government reports, web pages — not just journal articles.
- It runs a web search when CrossRef comes up empty instead of stopping at a database miss.
- It does not use Google Scholar. I wrote about why at some length: Scholar indexes references it sees in papers, so a hallucinated citation can end up minting its own Scholar record and then confirming itself.
The first bullet I suspect will cause the overall prevalence estimates in the prior large scale checks to be large underestimates. The later bullets could make the estimates be too high, but I believe not checking authors and journals likely dwarfs the other two errors.
To test this hypothesis, I took the 10 most recent articles posted to the hall of hallucinations, which contained 46 flagged hallucinations between them. I went through each flagged hallucination citation myself and assigned the error type. The breakdown:
- 14 wrong author (30%)
- 3 wrong journal only (7%)
- 29 wrong title or whole-cloth fabrication (63%)
Most of the papers contain a mix of errors, only one had 15 errors for totally whole cloth made up articles. The rest were portmanteaus of real authors/titles/journals, just Mad Libbed together into articles that do not actually exist.
So in the more typical situation — a paper with a handful of bad references rather than an entirely fabricated bibliography — close to half the errors are a real paper with the wrong people attached to it.
This is a small sample, so it may be non title errors are somewhere between 20% to 50%. So it is quite possible that the large scale estimates of AI hallucinated citations need to be doubled! (But to be more conservative should be increased by ~33%.)
If someone wants to help fund a large scale study, get in touch!
The papers classified
The 10 most recent hall of hallucinations entries at the time of writing, with the error breakdown I assigned to each:
- Report — 11 hallucinations: 6 wrong author, 2 wrong journal, 3 off in the title
- Report — 1 hallucination: 1 wrong title/book
- Report — 4 hallucinations: 2 wrong author, 1 wrong journal, 1 made up
- Report — 2 hallucinations: 2 titles
- Report — 6 hallucinations: 2 wrong author, 4 titles
- Report — 2 hallucinations: 1 wrong author, 1 title
- Report — 1 hallucination: 1 title
- Report — 2 hallucinations: 2 wrong author
- Report — 2 hallucinations: 1 wrong author, 1 wrong title
- Report — 15 hallucinations: all 15 title/whole cloth