Live data from Hacker News

Fake scientific papers are alarmingly common

science.org

81–90 of 105 posts

Re: Fake scientific papers are alarmingly common

#81
post #52

Earlier quoted context omitted.

My favourite example is the grievance studies ( https://en.wikipedia.org/wiki/Grievance_studies_affair ). The authors published, among others, portions of directly copied Mein Kampf. In one "study" submitted, they claimed to have observed thousands of hours of dogs having sex in parks, to observe the "patriarchal" linked to "rape culture". The entire thing is a horrible indictment on the level of scrutiny undertaken…

I have no sympathy for social science journals, but when you look at the details of this study, it's way less obvious that the rumor says it is. On the tens of articles they have proposed, the majority was rejected. The article "copied from Mein Kampf" was taking sentence from the book but changing words to create sentences that were scientifically correct (for example: "this social class is bad and we should avoid i…

Similarly, I read the one about dog parks, and if you approach it with the good faith notion that it wasn't just fabricated whole cloth, it reads like an okay-ish Master's level paper being submitted to an okay-ish journal. And features a decent sample size, which is a rare thing in the veterinary literature.

Re: Fake scientific papers are alarmingly common

#82
post #63
post #30

Earlier quoted context omitted.

No. They use 400 known fakes and 400 matched (presumed) non-fakes to estimate the sensitivity and specificity of their indicator, then apply that indicator to the full universe, then employ the estimated sensitivity and specificity to the obtained measurement to estimate the approximate actual rate of false papers. If you know the true prevalence of a disease in a population, and the sensitivity and specificity of yo…

I've worked on research estimating prevalence from imperfect tests, and something that concerns me about this study is that they aren't showing the error bars for their estimates. Typically, you would report a confidence interval for prevalence rather than just a point estimate, and the confidence intervals can often be fairly wide. There's two sources of uncertainty here, the assumed probabilistic nature of the diag…

Amazing. Reading more carefully, as FabHK pointed out above, they aren't even applying the obvious correction. They're just reporting the positive rate of the imperfect test. I've implemented Diggle's method [0]. When I have time, I'll see if they've provided enough data to do a proper analysis, and maybe write a blog post about it or something.

[0] https://github.com/indralab/opaque/blob/761572ed1b0d601271f0...

Re: Fake scientific papers are alarmingly common

#83
post #57

Earlier quoted context omitted.

> Furthermore, they’re explicitly saying that “red flagging” by their simple indicator doesn’t mean that the paper is fake, but that it merits higher scrutiny. Then they and science should change their sensationalist headline. It's ironic that a paper about fakeness of something uses a borderline misleading title.

You’re not wrong, but it is everyone’s own responsibility to read the article and not just the headline.

Expecting people to read every single article posted to HN is unrealistic.

Simply reading a title and on a topic you don’t find interesting then gives people the wrong impression.

Re: Fake scientific papers are alarmingly common

#84
post #57

Earlier quoted context omitted.

You’re not wrong, but it is everyone’s own responsibility to read the article and not just the headline.

So it's ok to lie in a portion of your work? Where do you draw the line? I draw it when someone starts communicating. Being wrong is ok, being deceitful isn't.

Is this headline really deceitful though? Certainly the research is flawed, but the statement "[bad thing] is alarmingly common" is basically just a subjective statement that lets you know what position the author is going to argue.

Re: Fake scientific papers are alarmingly common

#85
post #30

The researchers in this paper use an astonishingly biased "fake paper detector", requiring only two conditions to be met for any paper to be considered "fake": 1. Use a non-institutional email address, or have a hospital affiliation, 2. Have no international co-authors. And they acknowledge 86% sensitivity and 44% specificity. It's a coin-toss which biases massively against research from outside the US and Western Eu…

No. They use 400 known fakes and 400 matched (presumed) non-fakes to estimate the sensitivity and specificity of their indicator, then apply that indicator to the full universe, then employ the estimated sensitivity and specificity to the obtained measurement to estimate the approximate actual rate of false papers. If you know the true prevalence of a disease in a population, and the sensitivity and specificity of yo…

You can’t directly calculate both sensitivity and specificity using equal numbers of positives and negatives groups unless the actual population has that ratio.

A completely random test given equal populations results in 50% accuracy and 50% specificity. Things don’t look nearly as good if only 1% of the actual population has the condition.

Re: Fake scientific papers are alarmingly common

#86
post #48

Earlier quoted context omitted.

You know what's funny? Even if the numbers are hot garbage, they proved the point about how easy it is to publish fake science papers, since it got published. Kinda similar to those researchers years back who proved how easy it was to go into certain social science journals as long as you copied their ideology.

Well, there is a difference between "fake science" and "tried to do correct science but ending being wrong". If the second is "fake science", then basically all that Newton has ever produced is "fake science". For the social science journals bit, are you thinking of the "grievance studies affair": https://en.wikipedia.org/wiki/Grievance_studies_affair ? Ironically, this study has generated a lot of "fake news" on the…

There's also a difference between outright fake science i.e. lies/fabricated data in the manuscript and bad science i.e. the conclusions drawn by the authors were always "fake" because of bad practices but if you look at the details of the work they are honest about what they did. Of course ideally you would minimize both types of bad paper, but the latter isn't too damaging to the system in isolation while the former can cause a handful of papers to mislead a subfield of science for years. Also how to screen for and how to systemically discourage these two things could be quite different.

Re: Fake scientific papers are alarmingly common

#87
post #4

Unfortunately, it's an open secret that fake or low-effort almost useless papers are very common in every area of scientific research. Typically, it doesn't affect people working in that specific area - they develop/have a sixth sense to detect bullshit papers - it comes with experience but depends on several factors including the authors reputation, their institution (for the first screening), what journal/conferenc…

Right- academic papers are written for academicians who don't have any issues separating good papers and journals from bad. The fact that many journals have set themselves up as or allowed themselves to become part of the tenure and promotion metrics game, is more of an issue with tenure and promotion. If the requirement for simple metrics dissapeared, the fake papers would go away on their own. In any event, it's no…

I mean it is a problem for researchers though. The blatantly fake paper mill ones (which seem to be the topic of this article anyway) aren't, but scientific fraud or even just minor misconduct from people that know how to mask it can waste a great deal of grant funding and scientist time to figure out.

Like look how many times that 2006 Nature paper on amyloid beta in Alzheimer's was cited, turns out some of the images were completely fabricated.

Re: Fake scientific papers are alarmingly common

#88

Earlier quoted context omitted.

I agree that zero trust is in most cases a problematic goal. It's really root-of-trust vs web-of-trust that I'm on about here. If peer review is the product then the trust should be peer to peer. It feels like we're treating the publishers themselves as an authority, which I dislike.

Thank you for clarifying. The publishers ostensibly occupy a role of stewardship, I suspect the model must have made sense at one point. I admit its hard to see them as much more than rent extractors these days. The nature of trust relationships seems to trend towards aggregation and centralization. Do you have any thoughts on how a web of trust can sustain itself, or is that perhaps not a concern if a centralization…

There's a belief among some distributed systems folk:

> If your system doesn't have an explicit hierarchy then it has an implicit one.

I think it's hogwash. There are plenty of distributed systems in nature that lack a hierarchy (mycorrhizal networks in the soil of a forest come to mind). Truly distributed systems are possible, we humans are just bad at it.

Or rather, we're bad at designing for it. We do it all the time in our personal lives, we've been doing it for thousands of years, but when we introduce systems that are designed to scale globally, it falls apart and you end up with gatekeepers and renteeism.

Another distributed systems thing: the CAP theorem:

> Consistent, Availability, Partition Tolerance... chose two.

Usually, the systems we design are at the expense of partition tolerance (blockchains, for instance, go to great lengths to assure consistency).

But those fungal networks that I mentioned, they put partition tolerance first, which gives the system a sense of locality that is lacking when you instead focus on consistency.

That same sense of locality is found in natural emergent human social networks, they don't even try to achieve global consistency: if you think Jimbob is an asshat, and your friends agree, that's enough.

So I think the key to sustainable webs of trust lies somewhere in that underexplored design space where we make partition tolerance primary. Rather than building tech to tell people who to trust (think of that padlock icon in your browser) we should respect their autonomy a bit more and make the user experience be a function of that user's explicitly defined trust settings.

One thing I like about this is that it removes the edgelord dynamic. There's no advantage to being the guy who posts the most outrageous stuff that just barely squeaks by the moderator. Instead, everybody can publish, but if you want to be heard as widely as possible you need to be trusted (in whatever domain you're publishing in) by people who are themselves well trusted in that domain.

Experts can be found not by listening to some authority that tells you who the experts are, but instead by following the directed graph of trust relationships until you find a cycle. That cycle is a community of experts in the "trust color" you're querying for. So expertise is more emergent and less top-down.

If you can't agree with somebody about a topic, you can follow this graph and either find a mediator (someone you both transitively trust) or find separate experts who presumably exemplify the disagreement more energetically than you do.

Instead of:

> Agree with us or be silent

It would be more of a:

> Here's how we can disagree as fruitfully as possible

Navigating the resulting dataset and deciding what to believe would be left as an exercise to the user, which it already is, but we'd hopefully have given them enough so that we can scale further than our unaugmented trust instincts allow for.

There's unfortunately not much money in building things like this. There's no guarantee that you stay in control of what you've built (the users could just revoke trust in you while still using the software that you wrote) and that tends to be a turn-off for investors.

Re: Fake scientific papers are alarmingly common

#89
We must always put this in context, and I think we need to be careful about the narratives. Here's a few rules of thumbs

- Realistically the only people who can determine if a work is sound or not are other researchers in that same field.

- Peer review is a weak signal: reviewers are good at recognizing bad papers but not good at recognizing good papers (read this carefully).

- Most papers aren't highly influential. Thus meaning that we don't rely heavily on the results of most works (we rely weakly or purely for citations).

- The more influential a work is the more likely it is to be reproduced and scrutinized.

- Benchmarks are benchmarks, nothing more. Benchmarks are weak signals at best and shouldn't be used to make strong conclusions. Be that a p-value, FID, or even likelihood.

So we have to keep this in mind for a lot of reasons. One is how we discuss with the public. Headlines like this often make people grow wary of science. While scrutiny is good we have a good history of being successful. All processes are noisy but the cream has is more likely to come to the top and the surface is less noisy. It also tells us about who we should be listening to when taking advice and summaries of works. If you believe the news has failed us, then look to the sources.

I see many who only get their science from news sources that claim scientists are corrupt. I found this odd, especially considering I've worked at national labs and I can tell you that no one there is doing it for the money. You'd have to be a fucking idiot to do science for money. It doesn't pay well, you never get real time off, there is a high barrier to entry, and you are under high amounts of pressure. We're on a forum with Silicon Valley wages: the average physicist wage is 100k, what you'd make with a BS in CS but need an advanced degree for working at a lab. Let try to compare likes and likes by looking at LLNL. As a PhD physicist you'll make between $150k and $200/yr. You'll make the same as a PhD computer scientist. Yeah, this seems good, but we need to consider that if you drove 45 minutes west then that would be your base salary and you'd be making the same in other compensations. You can easily verify this and there's plenty of people you can ask for personal experience (I've seen people jump ship often). This doesn't prove that they aren't corrupt, but it provides strong evidence that if these people were motivated by monetary compensations (or even prestige) then there are far better opportunities for them.

Another important aspect, which I think is critical to forums like this, is to be careful how you as a non domain expert. Opinions are fine and no one should prevent you from having them. But the confidence in your opinion should be proportional to your qualifications. If you're an expert in one domain I'm sure you're frustrated by how many people discuss your domain as if they knew so much and they get so much wrong. How wrong answers float to the top of forums (HN and Reddit) and the gems are hidden. This usually comes down to a lack of nuanced understanding. Simple answers are almost never correct. Murry Gell-Mann amnesia doesn't just apply to reading the news. Discussions can be had without teaching. Scientific discussions aren't done through debate. Determine your goals, and ask yourself if the way you are discussing allows you to change your opinion or not. Make sure you're on the same page as others, using the same assumptions (this is a key failure point). I'll argue to go in with care. If you don't, you're just adding to the noise.

Re: Fake scientific papers are alarmingly common

#90
post #63
post #30

Earlier quoted context omitted.

No. They use 400 known fakes and 400 matched (presumed) non-fakes to estimate the sensitivity and specificity of their indicator, then apply that indicator to the full universe, then employ the estimated sensitivity and specificity to the obtained measurement to estimate the approximate actual rate of false papers. If you know the true prevalence of a disease in a population, and the sensitivity and specificity of yo…

I've worked on research estimating prevalence from imperfect tests, and something that concerns me about this study is that they aren't showing the error bars for their estimates. Typically, you would report a confidence interval for prevalence rather than just a point estimate, and the confidence intervals can often be fairly wide. There's two sources of uncertainty here, the assumed probabilistic nature of the diag…

> they aren't showing the error bars

Perhaps any paper without error bars should be tagged as a fake paper.

This one would have sneaked past though: https://retractionwatch.com/2022/12/05/a-paper-used-capital-...

Post reply on HN