Live data from Hacker News

AI intensifies fight against ‘paper mills’ that churn out fake research

nature.com

91–100 of 186 posts

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#91
post #14

Instead of fighting against "paper mills", let’s fight against journals. There are strong arguments against the need for peer-review for example [1]. Science does not get worse when there are more bad papers, science gets better when there are more good papers. Most papers in AI aren’t even reviewed by peers or editor and guess what: we have lots of progress happening. I’m not saying that is because the lack of revie…

Re: "Science does not get worse when there are more bad papers"... I don't think that this is true at all. Weeding through bad papers is, at a minimum, an opportunity cost, as is a good paper built on top of a bad one. Also, there is a societal cost in that bad research can get picked up and believed by people, like the anti-vax crowd. Or, bad research can be used to push an agenda, like anti-climate change.

> Also, there is a societal cost in that bad research can get picked up and believed by people, like the anti-vax crowd. Or, bad research can be used to push an agenda, like anti-climate change.

Responsible media would act as a gatekeeper. The problem is, most media utterly gutted scientific journalists for more profit, so they completely lack the basis to evaluate and supply context on research. Others, particularly boulevard media, willfully ignore any kind of ethics for clicks.

On top of that comes a general media illiteracy and media distrust that makes it even harder to combat because there's an awful lot of media that thrives on intentionally pushing crap to people.

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#92

Earlier quoted context omitted.

Imagine a world where all the raw data sits in a public repository next to a script to process the data which was reviewed as part of the publication process, which would accompany the paper and could be quickly and easily replicated with new data. Anyone could produce the graphs shown in the paper simply by downloading both and running them together. What a wonderful thing that would be. And you could imagine the go…

I hate to be sincere, but the reality is the data is our product. If we open source our data too much, we won't have anything left to publish as others who would have to expend nowhere near the resources we must do to produce it can scrape it and publish it. (that literally is how "AI" of the current hype bubble works today, lol, why would I want that online?) The incentives are definitely bad, and that's where the a…

If any level of government funds your institution you should have to releases the data.

The code I write is not my code, it's the banks.

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#93
post #14

Instead of fighting against "paper mills", let’s fight against journals. There are strong arguments against the need for peer-review for example [1]. Science does not get worse when there are more bad papers, science gets better when there are more good papers. Most papers in AI aren’t even reviewed by peers or editor and guess what: we have lots of progress happening. I’m not saying that is because the lack of revie…

This line of thinking is dangerous and ignores very real problems as the other comments point out. Instead of targeting peer-review, i recommend approaching this problem as one of incentives. Under the current system, what is the incentive for a reviewer to read and critique a paper? None. If they were paid in cash, there is a clear incentive. The publishers do not want this overheard and the associated legal require…

The incentives of paid peer review would not be easy to get right. Especially if you want the compensation to be high enough that it would actually be an incentive to the reviewer.

Many universities limit the number of hours their faculty can commit to paid external activities. Paid peer review would be one of those activities, and it would have to compete against other activities. Such as consulting, which can be very lucrative in some fields.

So maybe you have to pay $10k peer review fee when you submit a paper. And then it gets rejected after the first round of reviews, because you aimed too high or the editor and the reviewers just didn't like the topic. You resubmit to another journal and pay another $10k. After a couple of additional rounds of reviews ($5k each), the journal seems to be interested in publishing the paper. But reviewer 2 wants you to cite some of their papers that are not really that relevant. Do you agree, or do you argue against and risk another round of reviews (another $5k)? Or maybe you get the feeling that reviewer 3 is stalling the process with superficial requirements, as they get easy money from the reviews.

Except that most academics can't afford that. American academics are massively overpaid by global standards, and peer review would naturally be outsourced to developing countries, where the expectations of pay are more reasonable. Many of the academics there are reputable, after all. Unfortunately the institutional culture is often more problematic, and corruption also tends to be more prevalent. Do we really want to let those institutions shape the practices of science worldwide?

One of the unfortunate features of capitalism is that being a middleman is more profitable than doing the actual work. Instead of a system where you pay $x for reviews, we could end up with one where you pay $1.1x to a middleman, who then pays $0.8x for the reviews. The middlemen get even richer than in the current system, because there is more money in publishing than there used to be.

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#94
post #35

This seems like an extension of the replication crisis. In many fields, most published research is already bogus. The idea that peer review is enough to ensure that research is valid is perhaps not scaling well. It would be great to have things like open data sharing. At least in astronomy, which I'm somewhat familiar with, it doesn't seem like we're that close. Most scientists cannot even reproduce their own results…

(PhD student in STEM here) I think most people have the wrong idea about what peer review is. My advisor teaches us to treat it as a first check, but it is not a guarantee of correct results. Most of the time I spend on research is actually trying to understand the literature and reproducing their results. If I can't do it, it probably means I don't understand enough about the work I'm reading, but there is also the…

Most people have the wrond idea about what publication is, too. For almost all papers, for almost all readers, it doesn't matter if results are correct. It's just part of the game of academia. If you need to build something based on a paper, and you need it to work correctly (so, outside of public policy or macroeconomics), you need to do the work yourself.

But if you just want to cite the work and write your own paper on top of it, go ahead.

Published science is like ChatGPT ;-)

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#95
The easiest solutions won't be appreciated -- chain of trust from older professors/academics who co-author(or otherwise sign off that it's legit), metrics based on historical outputs from university+lab, and i'm sure some will start looking at names.

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#97
In the short term, the solution may be reputation, distributed by keychain.

When author C submits a paper and they're unknown to the journal, the journal consults the keychain to see whom has vouched for C as a credible researcher. If authors A and B have vouched for them, and A and B are in similarly good standing, C is regarded as being in good stead and the editorial process can move forward as usual.

If C is later found to have maliciously faked data, it isn't just C who gets dinged on the keychain, so do A and B by extension, providing an enforcement mechanism.

This is, in effect, how it has worked for centuries. Editor D calls/writes to A and B to ask, "hey, there's this new author C with a provocative paper that's in your subject area, but I've never heard of them. Are they the real deal?" If A and B vouch for C, but C turns out to be a fraud, Editor D will take any future consultation with A and B with a large grain of salt.

It is possible that AI will soon be able to generate entirely-credible looking papers. What AI will struggle to do, until it starts funding research of its own, is to generate novel research that reflects whatever is actually true of nature. A reputation system is the last bulwark against entropy, if no automated tools can sort wheat from chaff.

Lest you think versions of these systems aren't already in place, try posting something to the arXiv as a new author..... https://info.arxiv.org/help/endorsement.html

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#98
post #94

Earlier quoted context omitted.

(PhD student in STEM here) I think most people have the wrong idea about what peer review is. My advisor teaches us to treat it as a first check, but it is not a guarantee of correct results. Most of the time I spend on research is actually trying to understand the literature and reproducing their results. If I can't do it, it probably means I don't understand enough about the work I'm reading, but there is also the…

Most people have the wrond idea about what publication is, too. For almost all papers, for almost all readers, it doesn't matter if results are correct. It's just part of the game of academia. If you need to build something based on a paper, and you need it to work correctly (so, outside of public policy or macroeconomics), you need to do the work yourself. But if you just want to cite the work and write your own pap…

scientific publication is like a JIRA ticket.

to get visibility of your work you need to close JIRA tickets, the more the better, since people look at aggregate metrics.

The more your JIRA tickets are quoted in company documentation the better.

but JIRA ticket is not code, in the same way publication is not science

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#99
post #35

This seems like an extension of the replication crisis. In many fields, most published research is already bogus. The idea that peer review is enough to ensure that research is valid is perhaps not scaling well. It would be great to have things like open data sharing. At least in astronomy, which I'm somewhat familiar with, it doesn't seem like we're that close. Most scientists cannot even reproduce their own results…

It's ironic and deeply shameful when whitepapers side-step the most fundamental aspect of the scientific method: reproducibility.

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#100
post #35

This seems like an extension of the replication crisis. In many fields, most published research is already bogus. The idea that peer review is enough to ensure that research is valid is perhaps not scaling well. It would be great to have things like open data sharing. At least in astronomy, which I'm somewhat familiar with, it doesn't seem like we're that close. Most scientists cannot even reproduce their own results…

It might help to expand on "bogus". Bogus has a few levels, going from "not good, but possible from a well-intentioned author trying to do the right thing" (low-level bogus) to "outright deception" (high-level bogus). Small sample sizes, statistical errors, and flawed (but honest) experiment design are all, I suggest, low-level bogus. Faking data and plagiarism are high-level bogus. I think peer review is capable of,…

It's also possible to be functionally bogus while doing everything 100% by the book. If you control things super well you can inadvertently make your result so narrow that it's basically meaningless, at least with the way the scientific system functions today. If people worked together in a more concrete way to build on prior results this effect might not be so bad.

One of my favorite examples is a mouse study that did not replicate between two genetically identical mouse populations raised in the same conditions run by the same lab. The difference was the supplier the two sets of mice originally came from, and the researchers were able to pin down the cause as differences in the gut microbiome between the two (and in fact one particular bacterial strain). That is an example of great research, but the vast majority of studies will never catch something like this before they publish because they will only use one mouse supplier as they keep things controlled while minimizing costs.

Because designed replication studies are fairly rare and people often do not officially publish when they find things that don't replicate in biology, we are approaching interpretation/downstream use of these highly controlled studies in an extremely inefficient way.

But that isn't really the fault of individual researchers. Technically they're applying the scientific method correctly to their niche problem. It's the definition of the problem coupled with how we combine results across groups that causes the inefficiencies. As problems of interest increase in complexity we can't define them on the scale of individual labs anymore, and for some reason we've addressed it by breaking them up into subproblems in this homogeneous way. Then we just assume that piecing together the little results will work great...

Anyways, I agree there is also straight up fraud and blatantly bad practices on the level of individual papers, it's definitely a continuum. Sometimes such bad results slip into the mainstream of science or even have a huge impact on subsequent funding directions like with the fabricated amyloid beta paper. But I do suspect that for the most part the blatantly bad work stays on the fringes, and the largest negative impact on scientific productivity actually comes from a level of abstraction up.

Post reply on HN