Most citations in papers are not directly related to the actual content, but are more vapid, tangentially related prior work. Because of this, some take identifying hallucinated citations as pearl clutching. Who cares? They do not tell us if the paper's actual empirical results are accurate or not.
A recent set of hallucinations my tool identified, by Microsoft researcher Badri N. Patro, really signals how AI-hallucinated citations can show such a lack of care in an article that it is hard to take the rest of the work seriously.
Specifically, Patro's arXiv preprint, Counting without numbers and finding without words, shows this pretty plainly. My tool identified citation 6 as a hallucination. The citation is below: an article in JAMA that I do not believe exists.
You may spot an issue with the citation, which includes an explanatory note right in it. Hurricane Katrina occurred in 2005, so it is not possible that an article in 2001 referenced a survey about it. Just reading the text of the article, most readers would never bother to identify this pretty blatant discrepancy.
Now, this does not directly impact any empirical results. (This paper does not have any that I can tell; it is a mere proposal.) Running Patro's other recent papers on arXiv through VerusCite shows similar errors:
- LLMOrbit: A Circular Taxonomy of Large Language Models — From Scaling Walls to Agentic AI Systems, 2026, https://arxiv.org/abs/2601.14053, V2, VerusCite report, 14 potential hallucinations
- NAKUL-Med: Spectral-Graph State Space Models with Dynamics Kernels for Medical Signals, https://arxiv.org/abs/2605.00871, VerusCite report, 1 potential hallucination
- Unsure but Certain: Uncovering the Representation-Confidence Gap in Diffusion Language Models, https://arxiv.org/abs/2608.08791, VerusCite report, 1 potential hallucination
- Beyond the Bidirectional Promise: Re-evaluating the Robustness of Diffusion Language Models, https://arxiv.org/abs/2607.27386, VerusCite report, 1 potential hallucination
So I cannot argue that 14 potential hallucinated references (in a paper with 171 references) establish without a doubt that the work is wrong in Patro's LLMOrbit paper. I think it is a pretty strong signal, though, that there are likely more fundamental errors in the empirical results in that paper.