Live data from Hacker News

AI intensifies fight against ‘paper mills’ that churn out fake research

nature.com

151–160 of 186 posts

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#151
post #146

Earlier quoted context omitted.

The height of trust rot is Science publishing Woo Suk Hwang's 2005 stem cell research. This was one of the top scientific journals publishing research that appeared to be a nobel-prize track line of research that would have been a breakthrough in medical treatment. Instead it tainted a line of research. The research results were fraudulent, claiming a much higher success rate at generating a stem cell line than what…

I wasn't aware of the this kind of thing in 2005. How do you think the trustworthiness of papers has been trending since then, generally speaking. It sounds like that was a bit of an outlier.

I don't think a single worse example has occurred since, even though there have been other high profile instances of falsifying data and other unethical conduct.

I think after the replication crisis the scientific community is more aware of the limitations and faults of the current system. This has impacted different fields in different ways, to varying degrees of improvements. Now, I think a lab like Hwang's would be meet with more skepticism from both readers and the top tier journals.

The biggest cause for skepticism now would be Hwang's claimed success rate with the techniques. If other labs couldn't reach similar levels of success then the most generous assumption would be that the techniques weren't described well enough in the articles. Compare this to CRISPR gene editing, which is a slightly more modern advancement in genetics that is valued because of how easily other labs can incorporate it.

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#152

The easiest solutions won't be appreciated -- chain of trust from older professors/academics who co-author(or otherwise sign off that it's legit), metrics based on historical outputs from university+lab, and i'm sure some will start looking at names.

That will help entrench rich/old/"well-established" labs and academics even more.

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#153
post #121

Earlier quoted context omitted.

How do you see this as scapegoating? The headline specifically says "intensifies", the article very clearly positions AI-fabricated data as an extension of existing problems, and I don't see anything in the article downplaying those existing problems (the entire closing section is about how the summit was on issues broader than AI).

Nature's reporting on the problem of paper mills is surprisingly high quality and honest, given that these reports directly attack the credibility of Nature itself (and many other journals).

Does the argument attack Nature and other journals or does it setup a position that makes journals even more important to provide filters, verification, etc. services?

What were facing is an explosion of Brandolini's law on a scale people are not prepared to deal with and when all financial incentives promote this direction and financial incentives rule the world, I'm not sure how we get around the issue in environments that support free speech.

I'm not proposing we curtail free speech but we have serious issues to deal with as a society in terms of the believability and sheer volume of false information. To some degree this has always been an issue but my concern is that there's a critical tipping point in a free speech environment where there's so much BS out there people completely stop believing any information not matter how reputable or valid it is.

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#154

From my perspective as an outsider, I have always been amazed by the use of papers in academic research as a means of communicating findings to the wider world. I find it problematic that these papers are often formatted in a way that makes them highly unreadable, with two columns and compressed text. In my opinion, adopting more modern methods of publishing research could greatly enhance the overall quality of resea…

you don't know how to read papers then. first read the abstract. then the conclusion/discussion. then the methodology to see if the study was conducted reasonably. papers are made to be read by scientists -- dumbing them down to be understandable by layman would waste time for little gain. that is the job of science communicators and journalists -- and even these often do a pretty poor job of it.

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#155

Earlier quoted context omitted.

No, papers are not always useful. That isn't an inherent property. They can be useful. It depends on the content of the paper. If you can't tell whether a paper is real or bullshit it doesn't matter if an AI or a human created it. Journalists will, heaven forbid, have to do real journalism instead of blindly trusting some piece of paper they found somewhere.

> If you can't tell whether a paper is real or bullshit it doesn't matter if an AI or a human created it. That's true, but now it's incredibly easy to create a human-quality bullshit paper. That wasn't true a year ago. It doesn't matter who created it, but it does matter how easy it was to create. > Journalists will, heaven forbid, have to do real journalism instead of blindly trusting some piece of paper they found…

> That's true, but now it's incredibly easy to create a human-quality bullshit paper.

No, it's incredibly easy to create human-quality syntax. Not meaning. The source of the value was never the paper or the syntax - it was the meaning of the paper. Whether a human or an AI creates bullshit matters not.

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#156

Earlier quoted context omitted.

Imagine a world where all the raw data sits in a public repository next to a script to process the data which was reviewed as part of the publication process, which would accompany the paper and could be quickly and easily replicated with new data. Anyone could produce the graphs shown in the paper simply by downloading both and running them together. What a wonderful thing that would be. And you could imagine the go…

I hate to be sincere, but the reality is the data is our product. If we open source our data too much, we won't have anything left to publish as others who would have to expend nowhere near the resources we must do to produce it can scrape it and publish it. (that literally is how "AI" of the current hype bubble works today, lol, why would I want that online?) The incentives are definitely bad, and that's where the a…

I don't think this is the main reason people don't open source data. If your data is high quality and others are using it to find interesting stuff then you are going to get citations and authorship on additional publications with minimal work, which sounds like a great ROI. The reason data doesn't get open sourced is because researchers are not actually confident in the data and don't want their other works to be exposed.

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#157
post #148
post #35

This seems like an extension of the replication crisis. In many fields, most published research is already bogus. The idea that peer review is enough to ensure that research is valid is perhaps not scaling well. It would be great to have things like open data sharing. At least in astronomy, which I'm somewhat familiar with, it doesn't seem like we're that close. Most scientists cannot even reproduce their own results…

> "The idea that peer review is enough to ensure that research is valid is perhaps not scaling well." As others have said, this is not what peer review has ever been for, at all. It only checks for gross omissions and violations of form and syntax that are obvious to other scientists who are in adjacent fields (not even necessarily the same one). It's a relict from the times when publications were not target metrics.…

I have had reviewers and editors dig throughly into work, including code, and it has never been met with anything other than appreciation. I've sent reviews in detailing what further test or figure might be needed to make a point more persuasive, or noting inconsistencies, and I encourage my students to do the same.

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#158
post #35

This seems like an extension of the replication crisis. In many fields, most published research is already bogus. The idea that peer review is enough to ensure that research is valid is perhaps not scaling well. It would be great to have things like open data sharing. At least in astronomy, which I'm somewhat familiar with, it doesn't seem like we're that close. Most scientists cannot even reproduce their own results…

(PhD student in STEM here) I think most people have the wrong idea about what peer review is. My advisor teaches us to treat it as a first check, but it is not a guarantee of correct results. Most of the time I spend on research is actually trying to understand the literature and reproducing their results. If I can't do it, it probably means I don't understand enough about the work I'm reading, but there is also the…

While you should never treat anything in the literature as having "a guarantee of correct results", it should be a careful first check.

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#159
post #78
post #65

Earlier quoted context omitted.

Funding the storing and serving of all of that data doesn't sound like a difficult problem to me. That has gotten SO cheap over the past couple of decades. There are plenty of well funded institutions that can support that kind of resource.

You'd be surprised. One headache there is what did you tell study participants you would do with the data? Did you say you'd keep it forever? Did you say 5 years? Who's in charge of making sure that this centralized repository isn't inappropriately holding and distributing data? Funding organizations also have different requirements. Then a script is only useful if paired with a set of libraries of a particular versi…

This.

"Your data is open and available in perpetuity for whatever use" is in deep conflict with how we think about human subjects data, often for very good reason.

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#160

Earlier quoted context omitted.

I hate to be sincere, but the reality is the data is our product. If we open source our data too much, we won't have anything left to publish as others who would have to expend nowhere near the resources we must do to produce it can scrape it and publish it. (that literally is how "AI" of the current hype bubble works today, lol, why would I want that online?) The incentives are definitely bad, and that's where the a…

I don't think this is the main reason people don't open source data. If your data is high quality and others are using it to find interesting stuff then you are going to get citations and authorship on additional publications with minimal work, which sounds like a great ROI. The reason data doesn't get open sourced is because researchers are not actually confident in the data and don't want their other works to be ex…

"If your data is high quality and others are using it to find interesting stuff then you are going to get citations and authorship on additional publications with minimal work..."

This has not been born out in my experience, or the experience of others. Data products are chronically undervalued and undercited in science, and do not come with guarantees of authorship unless you put up barriers to access them without it.

Open data is, at this point, a decision made for either ideological reasons or as a condition of funding and/or publication, not a decision that is itself usually "worth" the investment of curating a public data set.

Post reply on HN