Live data from Hacker News

AI intensifies fight against ‘paper mills’ that churn out fake research

nature.com

161–170 of 186 posts

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#161
post #157
post #148

Earlier quoted context omitted.

> "The idea that peer review is enough to ensure that research is valid is perhaps not scaling well." As others have said, this is not what peer review has ever been for, at all. It only checks for gross omissions and violations of form and syntax that are obvious to other scientists who are in adjacent fields (not even necessarily the same one). It's a relict from the times when publications were not target metrics.…

I have had reviewers and editors dig throughly into work, including code, and it has never been met with anything other than appreciation. I've sent reviews in detailing what further test or figure might be needed to make a point more persuasive, or noting inconsistencies, and I encourage my students to do the same.

> I have had reviewers and editors dig throughly into work, including code, and it has never been met with anything other than appreciation.

Yes obviously that kind of free and unexpectedly diligent labor would be appreciated.

> I've sent reviews in detailing what further test or figure might be needed to make a point more persuasive, or noting inconsistencies, and I encourage my students to do the same.

Yes these are the kinds of things peer reviewers do. Notably not replicating the work.

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#162
post #35

This seems like an extension of the replication crisis. In many fields, most published research is already bogus. The idea that peer review is enough to ensure that research is valid is perhaps not scaling well. It would be great to have things like open data sharing. At least in astronomy, which I'm somewhat familiar with, it doesn't seem like we're that close. Most scientists cannot even reproduce their own results…

I love the idea. But do be aware that the high-quality software engineering like this will require has a cost: be prepared for basic science costs to scale appropriately (or for productivity to be reduced.) More to the point, not all science can be usefully verified by unit tests. Even machine-verified mathematical proofs are entirely dependent on the definitions being correct, and that requires expert human analysis…

Given the non-modular budget for an NIH R01 has not been updated since Clinton was president, there is absolutely no preparation for basic science costs to scale appropriately with inflation, let alone adding something that is a lot of work for relatively minimal per-paper return.

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#163

Earlier quoted context omitted.

I hate to be sincere, but the reality is the data is our product. If we open source our data too much, we won't have anything left to publish as others who would have to expend nowhere near the resources we must do to produce it can scrape it and publish it. (that literally is how "AI" of the current hype bubble works today, lol, why would I want that online?) The incentives are definitely bad, and that's where the a…

If any level of government funds your institution you should have to releases the data. The code I write is not my code, it's the banks.

What if I use your medical records, which contain pretty trivial identifying information in them?

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#164
post #160

Earlier quoted context omitted.

I don't think this is the main reason people don't open source data. If your data is high quality and others are using it to find interesting stuff then you are going to get citations and authorship on additional publications with minimal work, which sounds like a great ROI. The reason data doesn't get open sourced is because researchers are not actually confident in the data and don't want their other works to be ex…

"If your data is high quality and others are using it to find interesting stuff then you are going to get citations and authorship on additional publications with minimal work..." This has not been born out in my experience, or the experience of others. Data products are chronically undervalued and undercited in science, and do not come with guarantees of authorship unless you put up barriers to access them without i…

This is the limit to what people do. In my field, even source code sometimes isn't free software. However, if you ask a researcher, they will give you a copy you can freely modify, because the act of asking means they know who you are as a person, and can later ask for citations when you publish. Data is similarly treated. Just putting it out there without any barriers guarantees that other parties will use it without citing you.

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#165
post #161
post #157

Earlier quoted context omitted.

I have had reviewers and editors dig throughly into work, including code, and it has never been met with anything other than appreciation. I've sent reviews in detailing what further test or figure might be needed to make a point more persuasive, or noting inconsistencies, and I encourage my students to do the same.

> I have had reviewers and editors dig throughly into work, including code, and it has never been met with anything other than appreciation. Yes obviously that kind of free and unexpectedly diligent labor would be appreciated. > I've sent reviews in detailing what further test or figure might be needed to make a point more persuasive, or noting inconsistencies, and I encourage my students to do the same. Yes these ar…

"Yes obviously that kind of free and unexpectedly diligent labor would be appreciated."

Then perhaps characterizing doing work - including often doing partial replication, independently deriving results, checking code to make sure it's both complete and consistent, etc. as "ruining your career and your advisor will also be mad" is unfair.

"I've sent reviews in detailing what further test or figure might be needed to make a point more persuasive, or noting inconsistencies, and I encourage my students to do the same."

Some of this is replication. "I cannot get from X to Y in your paper without the addition of Z" is, at its essence, a statement that a result cannot be replicated given what has been provided.

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#166
post #78
post #65

Earlier quoted context omitted.

Funding the storing and serving of all of that data doesn't sound like a difficult problem to me. That has gotten SO cheap over the past couple of decades. There are plenty of well funded institutions that can support that kind of resource.

You'd be surprised. One headache there is what did you tell study participants you would do with the data? Did you say you'd keep it forever? Did you say 5 years? Who's in charge of making sure that this centralized repository isn't inappropriately holding and distributing data? Funding organizations also have different requirements. Then a script is only useful if paired with a set of libraries of a particular versi…

This is an excellent comment, thank you.

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#167
post #35

This seems like an extension of the replication crisis. In many fields, most published research is already bogus. The idea that peer review is enough to ensure that research is valid is perhaps not scaling well. It would be great to have things like open data sharing. At least in astronomy, which I'm somewhat familiar with, it doesn't seem like we're that close. Most scientists cannot even reproduce their own results…

The current scientific system has long been known to have serious problems of incorrect results and conflicts of interest. This article seems like an attempt to pin the crisis on an AI scapegoat. From 2015, the Editor of The Lancet: The case against science is straightforward: much of the scientific literature, perhaps half, may simply be untrue. Afflicted by studies with small sample sizes, tiny effects, invalid exp…

I would note there is a difference between untrue and fraudulent.

"Afflicted by studies with small sample sizes, tiny effects..."

Neither these are inherently the result of malfeasance or fraud, or even poor statistical practice.

Sample size is a compromise between a number of things - statistical power, yes, but also trying to minimize the number of human or animal subjects involved, and to be frank, budget (as I've noted elsewhere, the NIH R01 non-modular budget hasn't changed since the 90's). Before there's a body of work done, statistical power is often speculative. What do we think the effect estimate will be. If we're lucky, maybe we have a mathematical model suggesting at least something. In that arena, it's likely that we may undershoot the needed sample size - though I'll note underpowered findings are null findings, and far less likely to get published (and likely hopeless in say...The Lancet). There are also some questions that we may still want answers to where the sample size is inherently small. There are a finite number of Ebola outbreaks, or veterinary clinics in the U.S. (both real examples).

Similarly, tiny effects are hard to estimate, but that doesn't mean they're bad to estimate. Something that increased the risk of death in American citizens by 1% for example, would have a relative risk of 1.01, which is as small as many medical and epidemiology journals are apt to report. Yet this would impact thousands of people. Measuring that may be very hard, and very noisy, producing a number of wrong answers, but it's not self-evidently a bad idea. Especially if we don't know if the effect is tiny ahead of time.

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#168
post #165
post #161

Earlier quoted context omitted.

> I have had reviewers and editors dig throughly into work, including code, and it has never been met with anything other than appreciation. Yes obviously that kind of free and unexpectedly diligent labor would be appreciated. > I've sent reviews in detailing what further test or figure might be needed to make a point more persuasive, or noting inconsistencies, and I encourage my students to do the same. Yes these ar…

"Yes obviously that kind of free and unexpectedly diligent labor would be appreciated." Then perhaps characterizing doing work - including often doing partial replication, independently deriving results, checking code to make sure it's both complete and consistent, etc. as "ruining your career and your advisor will also be mad" is unfair. "I've sent reviews in detailing what further test or figure might be needed to…

> "Then perhaps characterizing doing work - including often doing partial replication, independently deriving results, checking code to make sure it's both complete and consistent, etc. as "ruining your career and your advisor will also be made" is unfair."

You are mischaracterizing my argument in multiple ways. What I said was "Peer reviewers never replicate the work. If you are a grad student tapped for peer review and you spend the time to replicate the work, you are ruining your career and your advisor will also be mad." which I stand by. I don't think it's unfair. Maybe your reinterpretation is unfair, but I can't speak to that because you were the one who wrote it not me.

Then in response to you saying "I have had reviewers and editors dig throughly into work, including code, and it has never been met with anything other than appreciation." I said "Yes obviously that kind of free and unexpectedly diligent labor would be appreciated." which I also stand by.

In response to you saying "I've sent reviews in detailing what further test or figure might be needed to make a point more persuasive, or noting inconsistencies, and I encourage my students to do the same." I said "Yes these are the kinds of things peer reviewers do. Notably not replicating the work." which I also stand by. You are saying that "Some of this is replication." Well it's obviously not. Asking for further tests and figures and noticing inconsistencies are exactly the kind of things that peer reviewers usually do, and it's not replicating the work.

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#169
post #159
post #78

Earlier quoted context omitted.

You'd be surprised. One headache there is what did you tell study participants you would do with the data? Did you say you'd keep it forever? Did you say 5 years? Who's in charge of making sure that this centralized repository isn't inappropriately holding and distributing data? Funding organizations also have different requirements. Then a script is only useful if paired with a set of libraries of a particular versi…

This. "Your data is open and available in perpetuity for whatever use" is in deep conflict with how we think about human subjects data, often for very good reason.

One thing that was really jarring for me moving from the startup/consulting/advertising space to the research space was that human subjects data really gets deleted.

It's not just the deleted_at column being set, there's no backup, it's really gone. Every copy, forever.

I appreciate the ethics of it, and part of my reason for working in this area is because of these ethics, but even 5 years in there is so much reluctance to press that delete button.

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#170

There was an interesting claim I heard from Eric Weinstein regarding peer review that he characterized as a cancerous infiltration into all of science from some medical area? Aha, found it in my notes... """ 1:32:28 Eric: but let me jump in--peer review is a cancer from outer space. It came from the biomedical community, it invaded science. The old system, because I have to say this because many people who are now pr…

Well it's certainly in keeping with the view he has of himself as an under appreciated scientific genius, however I don't think it makes for a very compelling critique of peer-review. Frankly, all his boohoo-hooing about being shut out of the in-group at Harvard probably has more to do with him being an insufferable narcissist, rather than any attempt by the establishment to prevent heterodox views in physics from reaching the wider scientific community.
Post reply on HN