Beeminder-relevant study turns out to be fake, womp womp

Actually there are two studies and we just learned from Data Colada that Study 2 was fake (and in fact fails to replicate), with the analysis for Study 1 promised next. I don’t doubt it will be equally dispositive.

  • Study 1: Students with externally imposed, evenly spaced deadlines did better than students who could pick their own deadlines. (Also that students did pick earlier-than-necessary deadlines, but not earlier enough. So the idea was that students know they need commitment devices but need them more than they think.)
  • Study 2: For a paid proofreading task, it’s the same story: people perform best with evenly spaced deadlines, worst with just one deadline for everything at the end, and in between for using a commitment device to commit to their own deadlines.

So, yeah, Ariely is such a slimeball charlatan, it makes me furious.

Here’s me in a Beeminder Book Brigade (for Annie Duke’s Thinking In Bets) two years ago:

On Ariely, I admit to being fooled in 2008. I thought Predictably Irrational was great. I was pretty oblivious to the replication crisis back then [1] and thought it was a compelling collection of results in behavioral economics. It was certainly fun to read. Shortly after that (2009ish?) was when I learned he was dishonest, for reasons I can’t share publicly. In hindsight, I imagine the best parts of Predictably Irrational were describing other people’s research, not Ariely’s. Although, hmm, there’s at least the 2002 experiment where he showed that students were willing to use, and got value from using, a commitment device for homework deadlines. That still seems valuable. I don’t know.

I do think @zzq nails it with this:

Ariely has the feel of the showman con artist, showing off his cool “facts” about psychology without particularly caring if they are true or not.

But maybe that’s a slope he started slowly sliding down when he started writing books, starting with Predictably Irrational in 2008.

In any case, I think Annie Duke deserves very little blame for calling Ariely’s book excellent in 2018. The fact that she’s praising the book, not Ariely himself, matters, for one thing. Like @clivemeister says, the book might even still hold up, I’m not sure. It does predate the paper that exposed his fraud.

To be clear, I despise Ariely, and don’t want to sound like I’m defending him in any way, even though I guess I’ve done so a little bit here. I guess I’m doing so just enough to defend Annie Duke citing him.


[1] I think most of us were, though I know at least one person – the economist I learned the “boy who cried wolf” mnemonic from, in fact – who surprised me by admitting that there was a whole class of social science papers he simply dismissed as likely BS. It seemed kinda shockingly anti-science at the time but it succeeded in planting important seeds of doubt for me.

I guess the good news is that there are still many Ariely-unaffiliated studies showing that commitment devices work, some of which even replicate.

(And since one of Ariely’s coauthors sued Data Colada for 25 million dollars I guess I’ll carefully hedge here to say that all that’s shown is that the data for the commitment device study was either fabricated or severely tampered with and the metadata has “Last saved by Dan Ariely”. So whatever probability that implies on “Ariely faked it” is all I’m alleging. Plus the sky-high prior from the previous scandals. And of course I meant “alleged slimeball charlatan”. Alleged by me.)

4 Likes

Potential fraud aside, Study 1 is very obviously nonsense. It randomizes sections and pretends that’s the same as randomizing students. (“no random assignment of individuals to treatments”) You can’t do that. “t(97) = 3.03, p = 0.003”—that 97 number is utterly bogus, invented degrees of freedom that do not exist. Indeed, given 2 sections and 2 treatment conditions, there are zero degrees of freedom, and there is no valid t-test to be done. To even claim to use a t-test given their setup is a misunderstanding of what a t-test even is!

This reflects a fundamental lack of seriousness. What did the authors even think they were doing? Worse than that, how did anyone ever take this paper seriously in the first place? In my opinion, though fraud is certainly bad, the whole thing was just a cargo-cult of the scientific process in the first place. Does it even matter if they made up the data, given that all they did with that data was babble about it in a science-flavored way?

Ariely is a charlatan uninterested in the truth—and also, it seems, a fraudster. But he’s a charlatan even apart from that.

I see what you mean about randomizing the sections rather than the students. Any effect of deadline schemes is hopelessly confounded with, for example, time of day or any other difference between the sections. The authors noted (or so they claimed) that there at least wasn’t a difference in overall academic record between the two sections.

But I think the paper admitted all that. The one with the students was a warmup, pilot study. Study 2 was the one meant to be Proper Science.

At least I was taken in by it.

1 Like

Passing that 97 to the t-test is conceptually a type error, at least in a programming language which models such thing in the type system. So I don’t really buy the excuse that it’s a warmup pilot study—they weren’t actually doing science there, just pretending to. Someone who actually understood what a t-test conceptually is would not have made that mistake.

Calculating t(97) because that’s the number of students minus two is only slightly less insane than calculating t(2000) because that’s the year of publication minus two. You can’t just choose arbitrary vaguely associated numbers to plug onto the formula!

Yes, as you say, it’s also hopelessly confounded, etc. True. But my objection here is on a deeper level than that. Not just that the science is wrong, but that it’s nonesense, dressed in pseudoscientific jargon.

I don’t mean to criticize your not noticing it—the pseudoscience camouflages it, it’s easy not to notice. Rather, the point o am trying to make os that the issue goes deeper than the fraud, as bad as fraud is. I think that both this and the fraud reflect the same fundamental opposition to truth: a lack of understanding that the aim of science is to seek out truth.

1 Like

I don’t want to appear to be defending Ariely in any way but I think the case against him is more compelling when it doesn’t overextend itself. It’s damning enough even if we’re maximally charitable! And for maximum charity we could imagine that students were effectively randomized which section they were in. And apparently they were mostly remote students so maybe there were no time/location differences. Conceivably (not that this is likely, even before getting to the fraud part) it was effectively an RCT and the t-test was fine.

I mean, I applaud your high standards for rigor. Just that if this, counterfactually, were someone doing research in good faith and they made this error, I’d want to be way less harsh. Especially if they acknowledged the problem with not individually randomizing and proceeded to a more rigorous version of the study.

Maybe think of the t-test as establishing (again, if the data were real) a bound on how much the spaced-out deadlines could possibly have helped.

Or not a bound, technically, but Bayesian evidence? The spurious t-test tells us the treatment effect conditional on there being no confounding effects. (I think?) And it’s not crazy to suspect, a priori, minimal confounding effects. Point being, if we just treat it as suggestive and the motivation for a proper RCT, it seems unobjectionable. And the authors weren’t trying to be sneaky about this part so it feels like diluting the case against them to pillory them on it.

1 Like

You are right. I guess my point is that already long before all this about fraud (both this and the previous instance) I fully despised Ariely and knew him to be completely uninterested in the truth. The actual fraud allegations add very little on top of that.

I’m not saying plugging nonsense numbers into a t-test is the same as fraud, just that even without knowing anything about fraud allegations, based just on what’s in the public paper, he very clearly is peddling utter nonsense. And thus, in a way, the fraud doesn’t matter—the paper already was obvious nonsense.

It’s like hearing that the author of a paper alleging that he was visited by space aliens who gave him psychic powers had actually made the whole story up—I mean, sure, fraud is bad in and of itself, and maybe before the revelation you could speculate that he was merely deluded and incompetent, not malicious, but ultimately, it shouldn’t change the amount you trust the paper, which should have already been at zero.

1 Like

In 2021, Data Colada discovered fabricated data in a 2012 field study published in PNAS[5] by Lisa L. Shu, Nina Mazar, Francesca Gino, Dan Ariely, and Max H. Bazerman.[6][7] All of the study’s authors agreed with their assessment and the paper was retracted.[7] The authors also agreed that Ariely was the only author who had access to the data prior to transmitting it in its fraudulent form to Mazar, the analyst.

It’s not the first time when a study linked with Ariely has some fabricated data.

1 Like

Yes! I clearly think about Ariely too much because I had just blithely assumed everyone had that context. Thank you for linking that. I’ve also now properly blogged about it:

2 Likes

Do you have a blog post about some actual recommended reading around beeminder and akrasia? I learned about him from ChatGPT a few months ago.

Did you even run an experiment on aggregated beeminder data? You have ten years of independently reported data from different demographics that can tell us if self-imposed deadlines work.

1 Like

But no control group?

1 Like

Control group or no, you can get a gut feeling about people’s attitude from a few SQL queries.

The underlying thing that is really interesting to me is that how I (skorytnicki) compare to the other people on beeminder - am I more or less diligent, do I derail more or not etc.

2 Likes