Live data from Hacker News

Data detectives spotted fake numbers in a widely cited paper

economist.com

21–30 of 42 posts

Re: Data detectives spotted fake numbers in a widely cited paper

#21
post #2

It always suprises me how people don't do minimum effort for hiding this stuff. And no one notices. (or don't care about the obviously suspicious aspect) Do scientists actually read what they cite?

> Do scientists actually read what they cite? Yes, but few actually scrutinize the methodology of the studies. Statistics is really hard. It's easier to assume peer reviewers would have rejected the paper if it was bad.

Benford's law (https://en.wikipedia.org/wiki/Benford%27s_law) is a pretty well-known test for fraudulent numbers. Of course, it's not infallible and (depending on the nature of the fake) it may be possible to tailor the numbers to 'pass' it, but it's a good heuristic. I'd be curious to know whether it would have detected this - and, if so, whether it was indeed used.

Re: Data detectives spotted fake numbers in a widely cited paper

#24

A lot of science today is basically parallel construction. You start with a sexy story that you know will get you a lot of press, like "promising you will be honest actually makes you behave in an honest way" and then you just make that paper happen, however you can. Under the publish or perish system, scientists don't have time to actually research the topic, and imagine if it fails to confirm - you just wasted a lo…

I think it's important to bear in mind that even with the system working perfectly well, and perfect ethics, we should expect to see a lot of papers published with false results.

Lets say there are 200 propositions we want to test and are candidates for publication, that 20 of them are true and that our error rate is 5%. That means when we test the 20 that are true 19 of them will be accurately shown to be correct and 1 will be erroneously found false.

However when the other 180 propositions are tested 5% of them will be erroneously found to be true, that's 9 propositions. This means we will end up with 28 'successful' studies that make it into prestigious journals, about 1/3 of which are false positives.

And as I said, that's if the system works perfectly with no fraud whatsoever. Throw in some human error and it's not surprising if a fair few studies start to look pretty dodgy. Add in some genuine fraud too and you've got a full-on replication crisis with all the trimmings.

Re: Data detectives spotted fake numbers in a widely cited paper

#26

Earlier quoted context omitted.

There's a lot of value being abused in the term 'science'. Science is a highly valued concept but it's the result of following the scientific method, not the output of anyone with a postgrad.

Science is what scientists do, like politics is what politicians do? Either science can be critiqued as a social construct or it is an unimpeachable Platonic aspiration. I can see both perspectives. But, communicating that science itself is a somewhat messy social phenomena might be better as a long-term message for the public.

Politics is definitely not defined as what politicians do and science is, as the GP said, when somone follows the scientific method which is something that happens all the time, far from a Platonic aspiration.

Re: Data detectives spotted fake numbers in a widely cited paper

#27

A lot of science today is basically parallel construction. You start with a sexy story that you know will get you a lot of press, like "promising you will be honest actually makes you behave in an honest way" and then you just make that paper happen, however you can. Under the publish or perish system, scientists don't have time to actually research the topic, and imagine if it fails to confirm - you just wasted a lo…

I think it is important not to extrapolate from a replication crisis in one field (e.g. psychology) to all the other sciences, because the picture painted by this (to my knowledge) doesn't accurately describe the underlying practises.

Re: Data detectives spotted fake numbers in a widely cited paper

#28

I think the original blog post [0] or Andrew Gelman's discussion of it [1] are both better sources for technical details and some historical context. In particular, this is not the first such issue for Dan Ariely, as Gelman points out, he has a history of sketchy scientific ethics like doing media tours for studies that he knows failed to replicate. [0] http://datacolada.org/98 [1] https://statmodeling.stat.columbia.…

One amusing point is that much of Gelman's post has an error itself (that someone pointed out in the comments a couple weeks ago): the NPR interview was in 2017, so "Ariely, as a coauthor of this article, had to have known for at least half a year before the NPR story that this finding didn’t replicate." is incorrect. Maybe Gelman should retract that part :).

Annoyingly, the NPR transcript at [1] only has a small note "(SOUNDBITE OF ARCHIVED NPR BROADCAST)" at the top with no indication of when (the audio doesn't seem to have a date either). The podcast show notes are apparently the only recordation of the date. [2]

[1] https://www.npr.org/transcripts/805808486

[2] https://pbs.twimg.com/media/E9UQv8LWEAkpoR0?format=jpg&name=...

Re: Data detectives spotted fake numbers in a widely cited paper

#29

I think the original blog post [0] or Andrew Gelman's discussion of it [1] are both better sources for technical details and some historical context. In particular, this is not the first such issue for Dan Ariely, as Gelman points out, he has a history of sketchy scientific ethics like doing media tours for studies that he knows failed to replicate. [0] http://datacolada.org/98 [1] https://statmodeling.stat.columbia.…

The datacolada post makes the very reasonable request that all data should be released, and scientists should make that a standard thing to do by doing it themselves and requesting others do it. It feels like this could be applied retroactively too. In this case the 2012 authors still had the data that they released in 2020 which is how the analysis got done that showed evidence of fraud. Might be worth just asking a…

It’s not possible to release data in all circumstances. If you work with health data (I have worked with birth certificates, EMRs, inpatient discharge abstracts, drug prescription histories and other data) you can’t post it publicly. You have to promise not to include a table in the paper with a cell size of fewer than ten individuals!

For what it’s worth, the Trump administration attempted to make issuing new health and environmental regs harder by requiring public data disclosure. They did this entirely because they knew that much of the data could not be disclosed. So if you were studying, eg, the effects of some pollutant on a health outcome using private data, you wouldn’t be able to rely on that study in a regulatory context bc the data could not be published.

It’s a worthy idea, but there are exceptions for good reasons.

Re: Data detectives spotted fake numbers in a widely cited paper

#30
post #9

Earlier quoted context omitted.

Did YOU even check what you cite? The AUTHORS of the original paper got a dataset from a company. They didn't assume the fraud from the start and published the paper based on it. Later, when they tried to analyze the issue more in-depth, they couldn't replicate the results. THE ORIGINAL AUTHORS PUBLISHED a paper about a failure to replicate. It was just then that someone looked at the original data and found that it…

It's seems highly likely that Ariely did it, not the company.

Why do you say that?
Post reply on HN