I am shocked, shocked to learn that this 2002 paper by Dan Ariely is based on dubious data and doesn’t replicate.

8 min read Original article ↗

Dan Ariely—the much-decorated business school professor, retired Wall Street Journal columnist, Ted-talk star, NPR hero, Jeffrey Epstein contact, Founding Partner of Irrational Capital, insurance agent, inspiration for a hit TV sitcom, and author of the instant-classic children’s book, “The Adventures of Professor D.”—has had pretty much the worst professional luck that any scientist could have.

Over the period of decades, he keeps ending up as coauthor on journal articles don’t replicate and that turn out, through absolutely no fault of his own, to be based on dubious or fake data. Bad collaborators, lazy research assistants, missing computer files . . . who knows how this is all happening, but it’s bad luck for sure.

Uri Simonsohn, Joe Simmons, and Lief Nelson just came across another one:

A new paper in Psychological Science reports a failure to replicate Study 2 of Ariely and Wertenbroch’s influential article entitled, “Procrastination, Deadlines, and Performance: Self-Control by Precommitment.” The original study, published in Psychological Science in 2002, found that people performed better on a set of tasks when each task had its own externally imposed deadline than when people set their own deadlines or faced a single last-day deadline for all tasks. The paper has had a lasting influence. It has been assigned reading in many economics and psychology courses, and has more than 2,100 citations on Google Scholar. . . .

About 20 years ago, on April 20, 2006, one of the authors of the forthcoming replication, Kyle Hyndman, received the original data files in an email sent from [email protected] . . . .

But then when Hyndman and his coauthor, Aberto Bisin, attempted to reanalyze the original data as part of their replication effort, they report that this happened:

In October 2024, at the request of the editors, we shared with Dan Ariely an analysis of the contents from the file purportedly for their Study 2 and asked for permission to include a summary of it in the paper. Dan Ariely denied our request, arguing, among other things, that the files we received may not be the actual data. He did not subsequently provide us with any additional data from the original paper.

Damn. I hate when that happens.

Simonsohn, Simmons, and Nelson continue:

This motivated us to return to this paper and fully analyze the original data for the two main studies. We conclude that the data in Studies 1 and 2 were tampered with. . . .

Our assessment that the data were tampered with are based entirely on the analyses presented in our posts. Readers can review the evidence and draw their own conclusions. . . .

Now comes the fun stuff.

And, by “fun,” I mean “horrible.”

Here’s one of the graphs from the 2002 paper:

Wow—that looks pretty impressive! Everything’s exactly in order, the standard errors are small, but the gaps between conditions 1, 2, and 3 aren’t exactly equal. They’re slightly irregular: exactly equal could raise suspicion, but these data look like they could be real . . . at least they do, until you look at them more carefully.

Here come Simonsohn, Simmons, and Nelson to rain on the parade:

Red Flag #1: The Effect Is Too Big
As shown in the reprinted figure above, Ariely and Wertenbroch report a perfect pattern of results, for all three dependent variables, with a sample size of only 20 per condition. The effects are also large. Extremely, implausibly large. . . .

Red Flag #2: Duplicate Observations . . .
18 of the 20 participants in the Last Day Deadline condition had a “Corrections Twin”, another participant who found exactly the same number of errors for each of the three proofreading tasks. Interestingly, these twins have ID numbers that are exactly 10 positions apart (e.g., subject S1 and subject S11 are twins; so are S7 and S17; etc.). (There were no error twins in the other two conditions.)
The existence of so many of these twins – and all of them in only one condition – is inconsistent with these data being real. . . .

Red Flag #3: Things That Should Be Very Highly Correlated Aren’t Correlated At All
At the end of their study, Ariely and Wertenbroch purportedly “asked participants to evaluate their overall experience [of the proofreading task] on five attributes . . .
You might expect these judgments to be correlated. For example, if someone says they liked the task, you might also expect them to say that it was interesting.
In the replication, this was (super) true. Controlling for experimental condition, the partial correlation between liking and interest was, quite sensibly, close to perfect:

But in the original data, this relationship was not only imperfect; it was not there at all. Participants who said they liked the task more did not say that they found the task to be more interesting:

The problem is not limited to these subjective measures. Consider the fact that people did three very similar proofreading tasks, each with 100 mistakes. Surely, we’d expect people who do better on one task to also do better on another, nearly identical task. That simple fact should manifest in extremely large correlations between performance on one task and performance on another. And in the replication data it does, as the correlations range from +.74 to +.90. But in the original data it doesn’t, as the correlations range from +.03 to +.27. . . .

Red Flag #4: No Rounding In Self-Reported Minutes
As you’ll recall from a minute ago, Ariely and Wertenbroch (2002) purportedly asked participants to “estimate how much time they had spent on each of the three tasks” (p. 223). When people provide estimates like this, they tend to round. They usually say “20 minutes” or “30 minutes” instead of “17 minutes” or “32 minutes”. And, indeed, when the replicators asked people to report how many minutes they spent on each of the three tasks, 85% of them gave a round number:

This is what we’d expect humans to do.

But in the original data, they did not do that. Only 11.7% of estimated minutes were round, consistent with the 10% you’d expect by chance alone:

Simonsohn, Simmons, and Nelson conclude:

We are unable to generate a benign explanation for all of the anomalies presented here. The original findings are too large and yet they do not replicate; there are duplicated observations; correlations that should be very strong are often non-existent; and values that should be rounded are not rounded. Based on this evidence, we believe the data for Study 2 of Ariely and Wertenbroch (2002) were severely tampered with or fabricated to produce the desired results.

Can you believe the bad luck of Ariely, to have this happen to him over and over again? I imagine he’ll want to launch an investigation to catch the real faker or fakers. What a waste of time, though. This is effort that could otherwise be devoted to designing psychology experiments. Or delivering Ted talks. Or modifying paper shredders. Or something.

I recommend to Ariely that, moving forward, he choose his collaborators and research assistants more carefully, so that he doesn’t again get stuck with fabricated data.

Maybe he could have them sign some sort of honesty pledge?

In Ariely’s own words, Professor D is “a charming character doing his best, struggling, and using social science to his advantage.” That’s one way to put it!

P.S. I employ some humor in these posts, but this sort of story doesn’t make me happy, it makes me sad. It doesn’t surprise me—lots of people enjoy cheating, and some of them will find their way into academic research, and some of those people will be successful at it—indeed, cheating can make it easier to be successful, as you’re no longer bound by the truth—so I’m not surprised, but I hate to see it.

In all seriousness, this sort of fraud is disgraceful, and whoever did it should admit it and pay some restitution to the Association for Psychological Science and the thousands of subsequent researchers who have relied on these fake findings. They should also reimburse the government for funds received in any later research grants whose proposals relied on these results. Ariely’s a nice guy, I get that, but I think that at this point he should just share the names of his collaborators and research assistants who’ve been doing all this faking It’s not right that he gets all the blame and they get off scot-free.

P.P.S. See here for more from Simonsohn et al.

P.P.P.S. I also again want to thank Simonsohn et al. for their efforts here. They got sued for doing this sort of thing! But they’re not intimidated and they keep at it. Good for them.

I hope that, at the very least, Ariely can send a public note to Simonsohn et al. thanking them for uncovering all these problems in his published papers. When people point out problems in my published work, I appreciate it and I thank them.