Live data from Hacker News

AI intensifies fight against ‘paper mills’ that churn out fake research

nature.com

171–180 of 186 posts

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#171
post #66

Earlier quoted context omitted.

I understand how journals made sense during a time when distribution was a hard problem. Now with the internet I don’t get the reason for them to exist in the form it does in 2023.

A journal is a way of ensuring that the paper you are about to read meets a minimum standard of quality. My anecdote: the second worst paper I've ever read was an ArXiv paper, presented by a young researcher who had no idea of how bad this paper was, Had I not teared it down (which was a very unpleasant experience for everyone involved) it would have been further propagated by other young researchers. And the worst p…

[flagged]

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#172

Earlier quoted context omitted.

That's the point - it's an arbitrary cutoff, so the high frequency with which results cluster around that point is indicative of a problem. If a field accepts P=0.05 as significant it means 1 in 20 results can be false positives just by random chance. Now think about how many papers get published, and many of them report more than one thing. The right threshold should really be an order of magnitude lower. It's not O…

>If a field accepts P=0.05 as significant it means 1 in 20 results can be false positives just by random chance. That is not the correct interpretation of a p-value. See the ASA's statement on p-values.[0] >Researchers often wish to turn a p-value into a state- ment about the truth of a null hypothesis, or about the probability that random chance produced the observed data. The p-value is neither. It is a statement a…

Thanks for the ASA statement link.

Yes, obviously the whole notion of a hard threshold is a bit nonsensical to begin with, but as the ASA statement says "Pragmatic considerations often require binary, yes-no decisions". There doesn't seem any way around that. People face decisions like, shall we continue to fund this investigation? Yes/no. Should we recommend lifestyle changes to the public? Yes/no. It doesn't make sense to try and map a P value into a budget, for example. At some point you need a threshold (and likewise for effect size and other things). Of course at some level there is fuzzyness, which is why I said if a paper is mostly reporting 0.049 values then ... and I didn't specify what, exactly because the conclusion should be something like "fuzzily suspicious and should look closer".

So it's fine for the ASA to complain about statistical significance leading to "distortion of the scientific process", but their proposed alternative can be boiled down to doing more peer review, as most of the things they tell researchers to look at are things researchers will never conclude in the negative about their own results, like study design (because if they did they wouldn't have got to the point of calculating a P value to begin with).

Finally, I disagree when you say "That is false". Scientists aren't going to stop using statistical significance thresholds because they do need to make binary decisions at some point, and dropping the threshold they currently use by 10x would immediately yield major improvements in replicability and robustness.

Anyway, put P-hacking to one side if you dislike that discussion, it's fine. It's not actually the thing that bothers me the most when I read papers. A big gap between the prose summaries and what the data [analysis] actually shows is a much more common source of distortion, IMO.

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#173

Earlier quoted context omitted.

Nature's reporting on the problem of paper mills is surprisingly high quality and honest, given that these reports directly attack the credibility of Nature itself (and many other journals).

Does the argument attack Nature and other journals or does it setup a position that makes journals even more important to provide filters, verification, etc. services? What were facing is an explosion of Brandolini's law on a scale people are not prepared to deal with and when all financial incentives promote this direction and financial incentives rule the world, I'm not sure how we get around the issue in environme…

It could make journals even more important, but in practice the reason paper mills exist is because journals aren't doing much (visible?) QA on papers, so spamming them with fake claims and auto-generated papers is a viable business model.

Journals unfortunately aren't stepping up to meet the challenge. I was writing articles about this problem several years ago and things haven't noticeably improved. This thread is full of people pointing out the obvious solution - take claim audit seriously and start by paying people to do professional peer reviews. Their actual solutions tend to look like spam filters for paper submission queues. It's enough to be able to say they're doing something, but not enough to actually make a big difference especially post ChatGPT.

I'm not so pessimistic, I think people learn to discriminate between sources. A lot of people right now are learning to generalize to "experts aren't" but it's a rather more nuanced understanding than the media like to make out when you dig in, for instance people understand that "expert" in this context usually means public sector funded academic or civil servant, and not e.g. a roughneck on an oil rig, some UI programmer at Apple, the guy who fixed their car last weekend. They learn that the first form of expertise is the type where making false claims can be beneficial for the people making them, whereas if you lie a lot about oil on an oil rig eventually something will explode.

I think we'll eventually get to a place where incentives are better aligned, for instance where data collection and aggregation is fully divorced from the people analyzing it. A lot of the distortion in science comes from the fact that academics need to collect data but aren't rewarded much for doing it, so once a dataset is collected it gets kept secret and that allows for a lot of dishonest game playing. It also means people are strongly incentivized to see things in the data that aren't really there. If data collection and analysis was fully divorced, those issues would go away (you'd get other problems of course but would they be worse?).

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#174

Feels like AI makes the fight against paper mills easier. A group could submit AI-generated papers to different places and create a public index based on how successful they are at getting their nonsense papers accepted.

Reviewers are already overloaded with reviewing all kinds of papers and aren't paid for it. To increase their workload is an absolute worst outcome.

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#175

Earlier quoted context omitted.

Nature's reporting on the problem of paper mills is surprisingly high quality and honest, given that these reports directly attack the credibility of Nature itself (and many other journals).

Does the argument attack Nature and other journals or does it setup a position that makes journals even more important to provide filters, verification, etc. services? What were facing is an explosion of Brandolini's law on a scale people are not prepared to deal with and when all financial incentives promote this direction and financial incentives rule the world, I'm not sure how we get around the issue in environme…

In science it seems the adage "talk shit get hit" is coming home to roost.

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#176

Earlier quoted context omitted.

>If a field accepts P=0.05 as significant it means 1 in 20 results can be false positives just by random chance. That is not the correct interpretation of a p-value. See the ASA's statement on p-values.[0] >Researchers often wish to turn a p-value into a state- ment about the truth of a null hypothesis, or about the probability that random chance produced the observed data. The p-value is neither. It is a statement a…

Thanks for the ASA statement link. Yes, obviously the whole notion of a hard threshold is a bit nonsensical to begin with, but as the ASA statement says "Pragmatic considerations often require binary, yes-no decisions". There doesn't seem any way around that. People face decisions like, shall we continue to fund this investigation? Yes/no. Should we recommend lifestyle changes to the public? Yes/no. It doesn't make s…

>dropping the threshold they currently use by 10x would immediately yield major improvements in replicability and robustness

This is precisely the sort of misunderstanding that surrounds p-values, and it is what the statement aims to correct. Lowering the standard p-value threshold does not imply an improvement in these factors. The problems of p-hacking and creating false positives lie in the disclosure of the methodology, not the threshold of p-value. That is the whole point.

This will be my last comment in this chain. I'm just trying to clear up some persistent misconceptions that I see on the internet.

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#177

Earlier quoted context omitted.

I think peer review exists for reasons beyond building or affirming the reputation of a scholar. In particular, it’s regarded as a filter for eliminating the very worst-quality research. And judging from the manuscripts that came across my desk when I was still in academia, I’d say it’s performing this particular duty reasonably well. As such, the connection between AI fakery and peer-review isn’t all that clear cut…

I mostly agree with your assessment of peer review, but feel that it's gone too far. It's decades of metric hacking going unresolved. My research area is ML and it is definitely far out of hand. When I was in physics I'd probably be closer to your statement. But right now we have measurements that conclude "reviewers are good at identifying bad papers, but bad at identifying good papers." Either I think we need to ac…

Elimination of venue-based reviews is an interesting proposition. I’ll have to think on that.

The contrast between ML and physics is equally interesting, and seems to corroborate the experience of another poster, who expresses doubt about removing peer-review from long-haul studies in biological sciences. I’m still struggling to put a finger on why exactly this difference might exist, though.

Anyway, thanks for your thoughts :)

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#178

For 'lay' people here on HN, I just want to make a quick point: Peer-review does not mean that the reviewer is re-doing experiments or re-running analyses. They are only reviewing to see if the paper merits inclusion in the journal. Often this means telling the authors to do more experiments or check other things. But, to be clear, the peer-reviewer does not re-do things and check if they are 'right'

Although there is a type of peer review that includes redoing experiments and analysis: artifact evaluation. All of the top conferences in my field (real-time embedded systems) include this as an opt-in option, and papers get a special badge if they also pass artifact evaluation. I strongly believe that other fields in computer science would benefit by including and normalizing this process. Besides the reproducibili…

I'm nearly certain that your's is the only field that does anything like re-experimentation then. I'm in biotechy fields and it's a totally different beast out here man.

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#179

Earlier quoted context omitted.

I mostly agree with your assessment of peer review, but feel that it's gone too far. It's decades of metric hacking going unresolved. My research area is ML and it is definitely far out of hand. When I was in physics I'd probably be closer to your statement. But right now we have measurements that conclude "reviewers are good at identifying bad papers, but bad at identifying good papers." Either I think we need to ac…

Elimination of venue-based reviews is an interesting proposition. I’ll have to think on that. The contrast between ML and physics is equally interesting, and seems to corroborate the experience of another poster, who expresses doubt about removing peer-review from long-haul studies in biological sciences. I’m still struggling to put a finger on why exactly this difference might exist, though. Anyway, thanks for your…

The value of venues I see as facilitating networking and collaboration. Maybe these could be centered around invited speakers and presenters?

For the difference between ML and physics/other sciences, there are a few points that I see.

- First, ML submits to conferences rather than journals. This means typically you have one shot at acceptance, and possibly one chance to rebuttal. Other fields use journals, where there often is a push to accept and that the reviewer's job is to make the paper better rather than filter absolutely. This results in no one actually advocating for the paper to be published other than the authors (no incentive to accept). Worse, we have a zero sum game, so if reviewers are also authors of other works (common) they have an optimal strategy of rejecting other players' works (large incentive to reject).

- Second, ML is extremely hype driven. The field just moves lightning fast, it is insane. This leads to incremental improvements (technically true for all science) but muddies the waters since it is easy to reject on low novelty but if you don't release your work fast you get scooped and get rejected for again low novelty. Unfortunately we don't define concurrent work well (and I've had works actually rejected because we did not cite works that were publicly released (arxiv) around the same time as us or, worse, after we submitted to the conference!).

- Third, there are a large number of people submitting works (nearly 10k per conference), the field is wide (reducing domain expertise), and highly empirical (few learn theory and base understandings and instead rely on benchmark results to evaluate). The size brings with it a lot of noise and the empirical nature reduces rigor (noise in evaluation). The large number of works also results in low reviewer quality as chairs scramble to add reviewers to works (in my last review, 3 of 4 reviewers self reported 3/5 confidence scores, meaning they did not work in this domain). So we just have way more sources of noise than would be seen in many other fields. Consistency experiments support this as well: finding reviewers are good at rejecting bad works but bad at rejecting good works. Which concludes that reviewers over reject (default strategy).

I also appreciate your opinions. I certainly don't have all the answers. As I see it, we got to decide as a community. The only way to do that is to hear complaints, proposals, and challenges. We're all in this together so to me it is most important that we openly discuss these. Thank you for your thoughts and comments.

Re: AI intensifies fight against ‘paper mills’ that churn out fake research

#180

Earlier quoted context omitted.

On the other hand, maybe we'll abandon peer review and return to a pre 1950's like style? Everyone in ML just uses arxiv because the field is moving so fast. All work is realistically peer reviewed once it is out there. We treat conferences (more important than journals) as very noisy signals. Unfortunately, we still use those as metrics for completing degrees or hiring people. The way I would try to make this better…

I don't think ML is a good example, because it's an outlier within an outlier. CS is already a weird field, because it leans so heavily on conference papers that see a single round of reviews. Journal papers with multiple rounds of reviews are more common in other STEM fields. Because of this difference, peer review in CS is more focused on accepting/rejecting the paper than on improving it. And then the CS model bre…

Honestly, I agree with every word here. I'd rather see conferences as networking events than quality filters. Conferences should invite published works. The conference system just adds noise, and like you are suggesting, the single round results in a zero sum game and no advocates for the paper to be accepted/improved. Thus reviewers optimize their strategy: reject everything.

While I still have problems with peer review and still advocate for open publication, I do think many the the specific issues of ML could be resolved by focusing on journals rather than conferences.

Post reply on HN