Live data from Hacker News

“Is there a heuristic we might use to identify and flag questionable papers?”

statmodeling.stat.columbia.edu

41–50 of 109 posts

Re: “Is there a heuristic we might use to identify and flag questionable papers?”

#41
indeed there is no such heuristic. The question is deeper, it's about how to trust that scientists from a culture that is perceived to be alien are living up to familiar scientific standards. It's easier to identify fraud when it comes in a systematic way from a specific region of the world. However there is a lot of fraud and scientific-freeriding coming from western scientists too, it's just sporadic and not visible among the thousands of high-quality studies.

It's also wrong to call a paper 'questionable' without qualifier. Questionable can mean anything, and some of the most groundbreaking ideas were questionable when they were put in the wild. I suppose the author means technically flawed.

Re: “Is there a heuristic we might use to identify and flag questionable papers?”

#42

A good candidate could be Benford's law [1], which makes predictions about the distribution of digits. If the digits of the numbers found in the results section of the paper deviate too much from this law, it could be a red flag. [1] https://en.wikipedia.org/wiki/Benford%27s_law

Benford's "law" doesn't apply to everything, so it would depend a lot on the field and phenomenon being studied. And to be honest, once it was known to be used in this way, it would be simple to make fake data comply with the law.

True. But the "once it was known...make fake data to comply" issue applies to ~every fake-detecting method which could be applied.

Big picture, the goal is not to catch every fake. It's to (1) catch enough so that smarter fakers learn to submit their stuff elsewhere (vs. to the journal, conference, or institution in question, which really cares about fake detection). And (2) dimmer fakers are mostly caught ~promptly.

Re: “Is there a heuristic we might use to identify and flag questionable papers?”

#43
post #4

“When a measure becomes a target, it ceases to be a good measure” Hence, I'm afraid peer-review is still the only option left.

Peer review is OK but peer review is not curation. Everyone wants to get published , makes up a hypothesis that barely passes statistical significance, and gets published. Make up a hypothesis in your mind about anything , there is very high chance you will find a paper that supports it (and very low chance to find a paper that refutes it). This is like postmodernism: every possible idea is deemed to have equal weight, which leads to 1=2 . Sometimes it seems it 's more productive for a scientist to become a monk, stop communicataing and start working alone in her lab undistracted by the spam notifications of published science.

Re: “Is there a heuristic we might use to identify and flag questionable papers?”

#45

H1. If it's written in word, not good. H2. If there's no git repo with the complete code to reproduce all experiments, then you can assume it's bullshit.

Not a good heuristic, Word processors are analogous to GIMP in the photo editing world. It's what some people just know, even if Photoshop or Markdown/Pandoc/LaTeX math are just around the corner.

Re: “Is there a heuristic we might use to identify and flag questionable papers?”

#46

H1. If it's written in word, not good. H2. If there's no git repo with the complete code to reproduce all experiments, then you can assume it's bullshit.

Yup. “We did some experiments using models, here are the outputs” doesn’t cut it. I don’t trust you.

Re: “Is there a heuristic we might use to identify and flag questionable papers?”

#47
post #45

H1. If it's written in word, not good. H2. If there's no git repo with the complete code to reproduce all experiments, then you can assume it's bullshit.

Not a good heuristic, Word processors are analogous to GIMP in the photo editing world. It's what some people just know, even if Photoshop or Markdown/Pandoc/LaTeX math are just around the corner.

In math and computer science, papers witten in word are a definite mark of crackpotery. I agree it's arbitrary and nonsensical, but clearly a good heuristic.

Re: “Is there a heuristic we might use to identify and flag questionable papers?”

#49
post #28

Earlier quoted context omitted.

Which two countries? India and China?

Funny story...A few years ago, my online reference hunting turned up a paper from China with three Chinese authors, and also exactly the same paper (down to the formatting and footnotes) from three Indian authors in India.

It's entirely possible that these were both copies of another paper in another country.
Post reply on HN