Hallucinated Citations and Accuracy of Papers

VerusCite

2 min read Original article ↗

Most citations in papers are not directly related to the actual content, but are more vapid, tangentially related prior work. Because of this, some take identifying hallucinated citations as pearl clutching. Who cares? They do not tell us if the paper's actual empirical results are accurate or not.

A recent set of hallucinations my tool identified, by Microsoft researcher Badri N. Patro, really signals how AI-hallucinated citations can show such a lack of care in an article that it is hard to take the rest of the work seriously.

Specifically, Patro's arXiv preprint, Counting without numbers and finding without words, shows this pretty plainly. My tool identified citation 6 as a hallucination. The citation is below: an article in JAMA that I do not believe exists.

Bibliography entry 6 cites a 2001 JAMA article and adds a note about a Hurricane Katrina evacuation study
Citation 6 claims a 2001 JAMA article includes a Hurricane Katrina evacuation study.

You may spot an issue with the citation, which includes an explanatory note right in it. Hurricane Katrina occurred in 2005, so it is not possible that an article in 2001 referenced a survey about it. Just reading the text of the article, most readers would never bother to identify this pretty blatant discrepancy.

Paper text claims Hurricane Katrina revealed that 44 percent of people who refused evacuation could not bring their pets
The body text cites that 2001 article for a Hurricane Katrina evacuation statistic.

Now, this does not directly impact any empirical results. (This paper does not have any that I can tell; it is a mere proposal.) Running Patro's other recent papers on arXiv through VerusCite shows similar errors:

So I cannot argue that 14 potential hallucinated references (in a paper with 171 references) establish without a doubt that the work is wrong in Patro's LLMOrbit paper. I think it is a pretty strong signal, though, that there are likely more fundamental errors in the empirical results in that paper.