Live data from Hacker News

GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

gptzero.me

241–250 of 528 posts

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#241

Earlier quoted context omitted.

Not in most fields, unless misconduct is evident. (And what constitutes "misconduct" is cultural: if you have enough influence in a community, you can exert that influence on exactly where that definitional border lies.) Being wrong is not, and should not be , a career-ending move.

If we are aiming for quality, then being wrong absolutely should be. I would argue that is how it works in real life anyway. What we quibble over is what is the appropriate cutoff.

There's a big gulf between being wrong because you or a collaborator missed an uncontrolled confounding factor and falsifying or altering results. Science accepts that people sometimes make mistakes in their work because a) they can also be expected to miss something eventually and b) a lot of work is done by people in training in labs you're not directly in control of (collaborators). They already aim for quality and if you're consistently shown to be sloppy or incorrect when people try to use your work in their own.

The final bit is a thing I think most people miss when they think about replication. A lot of papers don't get replicated directly but their measurements do when other researchers try to use that data to perform their own experiments, at least in the more physical sciences this gets tougher the more human centric the research is. You can't fake or be wrong for long when you're writing papers about the properties of compounds and molecules. Someone is going to come try to base some new idea off your data and find out you're wrong when their experiment doesn't work. (or spend months trying to figure out what's wrong and finally double check the original data).

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#242
post #213

The innumeracy is load-bearing for the entire media ecosystem. If readers could do basic proportional reasoning, half of health journalism and most tech panic coverage would collapse overnight. GPTZero of course knows this. "100 hallucinations across 53 papers at prestigious conference" hits different than "0.07% of citations had issues, compared to unknown baseline, in papers whose actual findings remain valid."

I’m not sure that’s fair in this context. In the past, a single paper with questionable or falsified results at a top tier conference was big news. Something that casts doubt on the validity of 53 papers at a top AI conference is at least notable. > whose actual findings remain valid Remain valid according to who? The same group that missed hundreds of hallucinated citations?

Which of these papers had falsified results and not bad citations?

What is the base rate of bad citations pre-AI?

And finally yes. Peer review does not mean clicking every link in the footnotes to make sure the original paper didn't mislink, though I'm sure after this bruhaha this too will be automated.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#243

Earlier quoted context omitted.

> Eventually you may arrive at something like the H-index, which is defined as "The highest number H you can pick, where H is the number of papers you have written with H citations." It's the Google search algorithm all over again. And it's the certificate trust hierarchy all over again. We keep working on the same problems. Like the two cases I mentioned, this is a matter of making adjustments until you have the des…

Incentives. First X people that reproduce Y get Z percent of patent revenue. Or something similar.

Most papers generate zero patent revenue or even lead to patents at all. For major drugs maybe that works but we already have clinical trials before the drug goes to market that validate the efficacy of the drugs.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#244
It has been several years since the reviewing process for top AI conferences have been broken as hell, due to having too many submissions and only a few reviewers (up to the point that Masters students are reviewing the papers). It was only a matter of time before these conferences will be filled with AI-written papers.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#245
post #2

Yuck, this is going to really harm scientific research. There is already a problem with papers falsifying data/samples/etc, LLMs being able to put out plausible papers is just going to make it worse. On the bright side, maybe this will get the scientific community and science journalists to finally take reproducibility more seriously. I'd love to see future reporting that instead of saying "Research finds amazing che…

I heard that most papers in a given field are already not adding any value. (Maybe it depends on the field though.)

There seems to be a rule in every field that "99% of everything is crap." I guess AI adds a few more nines to the end of that.

The gems are lost in a sea of slop.

So I see useless output (e.g. crap on the app store) as having negative value, because it takes up time and space and energy that could have been spent on something good.

My point with all this is that it's not a new problem. It's always been about curation. But curation doesn't scale. It already didn't. I don't know what the answer to that looks like.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#246

Earlier quoted context omitted.

> I'd love to see future reporting that instead of saying "Research finds amazing chemical x which does y" you see "Researcher reproduces amazing results for chemical x which does y. First discovered by z". Most people (that I talk to, at least) in science agree that there's a reproducibility crisis. The challenge is there really isn't a good way to incentivize that work. Fundamentally (unless you're independent weal…

> The challenge is there really isn't a good way to incentivize that work. What if we got Undergrads (with hope of graduate studies) to do it? Could be a great way to train them on the skills required for research without the pressure of it also being novel?

Most interesting results are not so simple to recreate that would could reliably expect undergrads to do perform the replication even if we ignore the cost of the equipment and consumables that replication would need and the time/supervision required to walk them through the process.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#247
post #2

Yuck, this is going to really harm scientific research. There is already a problem with papers falsifying data/samples/etc, LLMs being able to put out plausible papers is just going to make it worse. On the bright side, maybe this will get the scientific community and science journalists to finally take reproducibility more seriously. I'd love to see future reporting that instead of saying "Research finds amazing che…

In my mental model, the fundamental problem of reproducibility is that scientists have very hard time to find a penny to fund such research. No one wants to grant “hey I need $1m and 2 years to validate the paper from last year which looks suspicious”. Until we can change how we fund science on the fundamental level; how we assign grants — it will be indeed very hard problem to deal with.

Funding is definitely a problem, but frankly reproduction is common. If you build off someone else's work (as is the norm) you need to reproduce first.

But without repetition being impactful to your career and the pressure to quickly and constantly push new work, a failure to reproduce is generally considered a reason to move on and tackle a different domain. It takes longer to trace the failure and the bar is higher to counter an existing work. It's much more likely you've made a subtle mistake. It's much more likely the other work had a subtle success. It's much more likely the other work simply wasn't written such that a work could be sufficiently reproduced.

I speak from experience too. I still remember in grad school I was failing to reproduce a work that was the main competitor to the work I had done (I needed to create comparisons). I emailed the author and got no response. Luckily my advisor knew the author's advisor and we got a meeting set up and I got the code. It didn't do what was claimed in the paper and the code structure wasn't what was described either. The result? My work didn't get published and we moved on. The other work was from a top 10 school and the choice was to burn a bridge and put a black mark on my reputation (from someone with far more merit and prestige) or move on.

That type of thing won't change in a reproduction system but needs an open system and open reproduction system as well. Mistakes are common and we shouldn't punish them. The only way to solve these issues is openness

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#248

Earlier quoted context omitted.

If you have a strong grip on exactly what it means, sure, but look at any HN thread on the topic of fraud in science. People think replication = validity because it's been described as the replication crisis for the last 15 years. And that's the best case! Funding replication studies in the current environment would just lead to lots of invalid papers being promoted as "fully replicated" and people would be fooled ev…

while i agree that "reproducibility is overrated", i went ahead and read your medium post. my feedback to you is, my summary of that writing: "mike_hearn's take on policy-adjacent writing conducted by public health officials and published in journals that interacted with mike_hearn's valid and common but nonetheless subjective political dispute about COVID-19." i don't know how any of that writing generalizes to othe…

Thanks for reading it, or scan reading it maybe. Of the 18 papers discussed in the essay here's what they're about in order:

- Alzheimers

- Cancer

- Alzheimers

- Skin lesions (first paper discussed in the linked blog post)

- Epidemiology (COVID)

- Epidemiology (COVID, foot and mouth disease, Zika)

- Misinformation/bot studies

- More misinformation/bot studies

- Archaeology/history

- PCR testing (in general, discussion opens with testing of whooping cough)

- Psychology, twice (assuming you count "men would like to be more muscular" as a psych claim)

- Misinformation studies

- COVID (the highlighted errors in the paper are objective, not subjective)

- COVID (the highlighted errors are software bugs, i.e. objective)

- COVID (a fake replication report that didn't successfully replicate anything)

- Public health (from 2010)

- Social science

Your summary of this as being about a "valid and common but subjective political dispute" I don't agree is accurate. There's no politics involved in any of these discussions or problems, just bad science.

Immunology has the same issues as most other medical fields. Sure, there's also fraud that requires genuinely deep expertise to find, but there's plenty that doesn't. Here's a random immunology paper from a few days ago identified as having image duplications, Photoshopping of western blots, numerous irrelevant citations and weird sentence breaks all suggestive that the paper might have been entirely faked or at least partly generated by AI: https://pubpeer.com/publications/FE6C57F66429DE2A9B88FD245DD...

The authors reply, claiming the problems are just rank incompetence, and each time someone finds yet another problem with the paper leading to yet another apology and proclamation of incompetence. It's just another day on PubPeer, nothing special about this paper. I plucked it off the front page. Zero wet lab experience is needed to understand why the exact same image being presented as two different things in two different papers is a problem.

And as for other fields, they're often extremely shallow. I actually am an expert in bot detection but that doesn't help at all in detecting validity errors in social science papers, because they do things like define a bot as anyone who tweets five times after midnight from a smartphone. A 10 year old could notice that this isn't true.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#249
post #232

This is going to be a huge problem for conferences. While journals have a longer time to get things right, as a conference reviewer (for IEEE conferences) I was often asked to review 20+ papers in a short time to determine who gets a full paper, who gets to present just a poster, etc. There was normally a second round, but often these would just look at submissions near the cutoff margin in the rankings. Obvious slop…

AI conferences are already fucked. Students who are doing their Master's degrees are reviewing those top-tier papers, since there are just too many submissions for existing reviewers.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#250
post #36

Earlier quoted context omitted.

AFAIK, no, but I could see there being cause to push citations to also cite the validations. It'd be good if standard practice turned into something like Paper A, by bob, bill, brad. Validated by Paper B by carol, clare, charlotte. or Paper A, by bob, bill, brad. Unvalidated.

Academics typically use citation count and popularity as a rough proxy for validation. It's certainly not perfect, but it is something that people think about. Semantic Scholar in particular is doing great work in this area, making it easy to see who cites who: https://www.semanticscholar.org/ Google Scholar's PDF reader extension turns every hyperlinked citation into a popout card that shows citation counts inline i…

That is a factor most people miss when thinking about the replication crisis. For the harder physical sciences a wrong paper will fairly quickly be found because as people go to expand on the ideas/use that data and get results that don't match the model informed by paper X they're going to eventually figure out that X is wrong. There might be issues with getting incentives to write and publish that negative result but each paper where the results of a previous paper are actually used in the new paper is a form of replication.
Post reply on HN