Live data from Hacker News

AI intensifies fight against ‘paper mills’ that churn out fake research

nature.com

101–110 of 186 posts

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#101
post #14

Instead of fighting against "paper mills", let’s fight against journals. There are strong arguments against the need for peer-review for example [1]. Science does not get worse when there are more bad papers, science gets better when there are more good papers. Most papers in AI aren’t even reviewed by peers or editor and guess what: we have lots of progress happening. I’m not saying that is because the lack of revie…

Re: "Science does not get worse when there are more bad papers"... I don't think that this is true at all. Weeding through bad papers is, at a minimum, an opportunity cost, as is a good paper built on top of a bad one. Also, there is a societal cost in that bad research can get picked up and believed by people, like the anti-vax crowd. Or, bad research can be used to push an agenda, like anti-climate change.

The solution to this is reputation. The flaw is that we're using the reputation of for-profit journals rather than the research institutions.

You shouldn't have confidence in a paper because it was published in Nature, you should have confidence in it because it was published by Harvard or the University of California or Google Research, who puts the name of their institution on it and thereby stakes their reputation. For the preeminent researchers in a field, their own names may be enough for people in that field to trust the result.

You can still have a paper published by nobody, and if the results are interesting, researchers at known institutions can try to replicate it. Until then it's nothing. But if they do, the nobody gets cited and credited with being the first and takes a step to building their own reputation, and the known institution gets credited with knowing when to spend resources following up on something significant.

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#102
post #35

This seems like an extension of the replication crisis. In many fields, most published research is already bogus. The idea that peer review is enough to ensure that research is valid is perhaps not scaling well. It would be great to have things like open data sharing. At least in astronomy, which I'm somewhat familiar with, it doesn't seem like we're that close. Most scientists cannot even reproduce their own results…

(PhD student in STEM here) I think most people have the wrong idea about what peer review is. My advisor teaches us to treat it as a first check, but it is not a guarantee of correct results. Most of the time I spend on research is actually trying to understand the literature and reproducing their results. If I can't do it, it probably means I don't understand enough about the work I'm reading, but there is also the…

It's never going to be a guarantee, but it's also implemented very shittily at present. There's a difference between what peer review is and what it should be. It makes sense to look at it realistically in day to day practice, but in discussion of systemic problems it should be viewed in a different light.

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#103

Earlier quoted context omitted.

Re: "Science does not get worse when there are more bad papers"... I don't think that this is true at all. Weeding through bad papers is, at a minimum, an opportunity cost, as is a good paper built on top of a bad one. Also, there is a societal cost in that bad research can get picked up and believed by people, like the anti-vax crowd. Or, bad research can be used to push an agenda, like anti-climate change.

You are correct. Many people want research to work like a social media site where everyone can contribute, believing this is an egalitarian--and therefore better--solution. In reality, it would slow research to a crawl, as noise and conflicted interests dominate the conversation and drown out more rigorous or valuable information.

This depends heavily on the field.

In many cases results can easily be independently verified. This is why it works for AI. If you publish a result, it should come with code anybody can run. If the code doesn't exist or doesn't do what you say it does, you're a fool and everyone can ignore you. If it does, you don't need anyone's stamp of approval to prove it.

But that doesn't work with medical trials or things of that nature where independently verifying the claims is expensive.

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#104

Earlier quoted context omitted.

I hate to be sincere, but the reality is the data is our product. If we open source our data too much, we won't have anything left to publish as others who would have to expend nowhere near the resources we must do to produce it can scrape it and publish it. (that literally is how "AI" of the current hype bubble works today, lol, why would I want that online?) The incentives are definitely bad, and that's where the a…

If any level of government funds your institution you should have to releases the data. The code I write is not my code, it's the banks.

If that happened it would indirectly change the incentives anyway, because everyone would start being required to release data, so bad practices that are currently incentivized would become impossible.

However I think it would be better to directly incentivize data release rather than require it, at least in biomedical sciences. Because of patient privacy issues there is no way raw data release can be required across the board. And I certainly do not trust the NIH to come up with a coherent set of rules for when it is versus isn't allowed, which would mean loopholes and more corruption.

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#105
post #8

Earlier quoted context omitted.

That doesn't really seem relevant to the discussion. AI obviously makes it 1000x easier to generate realistic-looking fake papers, so it "intensifies" the issue as this article explains. It is newsworthy and discussion-worthy, and as you said AI doesn't have agency so we don't need to worry about hurting its feelings by referencing it in a headline.

The only reason it's a problem is that papers became the measure - more papers = more prestige. The measure will simply change because now it won't have a strong signal.

That is not the only reason and it's not the most important reason. Papers are useful in the real world and people use them to learn stuff, and also reference them as evidence. If you're trying to learn about a topic and you have to wade through 90% bullshit AI papers when search for the topic, when you aren't experienced enough to immediately tell which papers are real, that's a problem. If malicious actors are able to spread disinformation that cites very real-looking research, or pitch journalists with very real-looking papers, that's a problem.

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#107

Earlier quoted context omitted.

This line of thinking is dangerous and ignores very real problems as the other comments point out. Instead of targeting peer-review, i recommend approaching this problem as one of incentives. Under the current system, what is the incentive for a reviewer to read and critique a paper? None. If they were paid in cash, there is a clear incentive. The publishers do not want this overheard and the associated legal require…

The incentives of paid peer review would not be easy to get right. Especially if you want the compensation to be high enough that it would actually be an incentive to the reviewer. Many universities limit the number of hours their faculty can commit to paid external activities. Paid peer review would be one of those activities, and it would have to compete against other activities. Such as consulting, which can be ve…

This makes a good point but is largely a strawman. Review fees do not need to be this high. At $ 25 an hour, most engineering papers take about 2-5 hours to review. That’s $ 125 per reviewer or $ 500 for three reviewers, one editor and another $ 500 for publication fees. That’s $ 1,000 in total and the number of reviewers can be fewer.

At the moment, greed drives open access publication fees. Journals charge thousands of dollars because they can. There is no logical justification of the cost. There are several journals that charge hundreds of euros for review although they do not pay reviewers.

Note also that not all review needs to be done by academics. folks in industry can contribute equally but their employer pays for their time so that boils down to altruism or business priorities. A reasonable payment for time creates an incentive for the world beyond academia and this is a net positive. The payment does not need to be prohibitive, just enough that the reviewer has an incentive to do a good job and isn’t pressured to rush. The problem thus is not limited to the journals or the peer review process but inventives.

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#108
post #35

This seems like an extension of the replication crisis. In many fields, most published research is already bogus. The idea that peer review is enough to ensure that research is valid is perhaps not scaling well. It would be great to have things like open data sharing. At least in astronomy, which I'm somewhat familiar with, it doesn't seem like we're that close. Most scientists cannot even reproduce their own results…

(PhD student in STEM here) I think most people have the wrong idea about what peer review is. My advisor teaches us to treat it as a first check, but it is not a guarantee of correct results. Most of the time I spend on research is actually trying to understand the literature and reproducing their results. If I can't do it, it probably means I don't understand enough about the work I'm reading, but there is also the…

Peer review cannot catch if they performed the procedure correctly (e.g. in a lab), but it is there to check things like:

- Validity of the experimental setup

- Validity of the statistical analysis methodology

- Validity of the conclusions

Take a look at the items in this comment:[1]

> Afflicted by studies with small sample sizes, tiny effects, invalid exploratory analyses

All these can be caught with peer review.

> together with an obsession for pursuing fashionable trends of dubious importance

In contrast, this problem is intensified with peer review.

> In their quest for telling a compelling story, scientists too often sculpt data to fit their preferred theory of the world. Or they retrofit hypotheses to fit their data.

Not the role of peer review to catch these, but making your data and analyses scripts open will allow these to be caught by anyone who wishes to. The problem we have with current publishing is that I could use this technique to get faulty results, publish in a prestigious journal, get a 100 citations on my paper, and one of those citations will be from the person who looked at my data and saw obvious biases in my selection (which cannot be caught by a mere peer review). That one citation pointing out how wrong I am is lost in the noise.

We need a way to highlight that citation - merely making a simple graph of connections via references doesn't get us there.

[1] https://news.ycombinator.com/item?id=36140540

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#109
post #22

Feels like AI makes the fight against paper mills easier. A group could submit AI-generated papers to different places and create a public index based on how successful they are at getting their nonsense papers accepted.

https://en.wikipedia.org/wiki/Sokal_affair

Sokal did not publish in a peer reviewed journal.

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#110

Prediction: as a result of fake research and it’s ilk (fake credentials, fake trade-journal pubs, fake reviews, etc.), we will observe a renewed interest in referrals and meatspace networking. It will become correspondingly difficult for outsiders to enter professional circles.

On the other hand, maybe we'll abandon peer review and return to a pre 1950's like style? Everyone in ML just uses arxiv because the field is moving so fast. All work is realistically peer reviewed once it is out there. We treat conferences (more important than journals) as very noisy signals. Unfortunately, we still use those as metrics for completing degrees or hiring people. The way I would try to make this better…

I think peer review exists for reasons beyond building or affirming the reputation of a scholar. In particular, it’s regarded as a filter for eliminating the very worst-quality research. And judging from the manuscripts that came across my desk when I was still in academia, I’d say it’s performing this particular duty reasonably well. As such, the connection between AI fakery and peer-review isn’t all that clear cut to me.

I might expect the peer-review pipeline to become inundated with prima facie credible manuscripts, which would overwhelm the reviewer’s ability to process publications even more than is already the case. At this point, I’d expect scholarly reputation to become even more important for getting published, and I’d expect to see an increase in signaling and other out-of-band communications (e.g. talking about your draft at a conference). Critically, I’d expect this to generally avoid AI generated “studies”, despite the obvious drawbacks.

And moreover, I actually have serious reservations about the project of purging the scientific landscape of informal processes. While there are some obvious and notorious drawbacks to the way things currently work, I am also wary of top-down, “authoritarian high-modernist” (to borrow the turn of phrase from James C. Scott’s Seeing Like a State) projects in all their forms, especially in science. The production of scientific knowledge is not pure technae; it cannot be reduced to a well-behaved formula. A large part of it is a craft requiring a significant share of informal processes in order to be truly creative. But that is a separate topic, and indeed, one might argue that doing away with formal peer-review altogether would be a step in the direction of informality.

So at the end of the day, I’m not convinced by any predictions that relate the effects of AI on peer-review. I think it could go either way :)

Post reply on HN