Live data from Hacker News

GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

gptzero.me

291–300 of 528 posts

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#291
post #279
post #273

Earlier quoted context omitted.

> I'd like to see the software engineers on this site say with a straight face that writing bugs should lead to jail time. My hand is up. I do not believe in gaol, but I do agree with the sentiment.

Let he who is without sin cast the first stone…

If there were real consequences, we wouldn't be forced to churn out buggy nonsense by our employers. So we'd be able to take the time to do the right thing. Bug free software is possible, the world just says its not worth it today.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#292
I searched Google for one of the hallucinations: [N. Flammarion. Chen "sam generalizes"]

AI Overview: Based on the research, [Chen and N. Flammarion (2022)](https://gptzero.me/news/neurips/) investigate why Sharpness-Aware Minimization (SAM) generalizes better than SGD, focusing on optimization perspectives

The link is a link to the OP web page calling the "research" a hallucination.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#293
post #238

I spot-checked one of the flagged papers (from Google, co-authored by a colleague of mine) The paper was https://openreview.net/forum?id=0ZnXGzLcOg and the problem flagged was "Two authors are omitted and one (Kyle Richardson) is added. This paper was published at ICLR 2024." I.e., for one cited paper, the author list was off and the venue was wrong. And this citation was mentioned in the background section of the pa…

I see your point, but I don’t see where the author makes any claims about the specifics of the hallucinations, or their impact on the papers’ broader validity. Indeed, I would have found the removal of supposed “innocuous” examples to be far more deceptive than simply calling a spade a spade, and allowing the data to speak for itself.

The point is that they should focus on the meaningful errors, not the automiation of meaningless errors.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#294
post #195
post #150

Earlier quoted context omitted.

I'd personally like to see top conferences grow a "reproducibility" track. Each submission would be a short tech report that chooses some other paper to re-implement. Cap 'em at three pages, have a lightweight review process. Maybe there could be artifacts (git repositories, etc) that accompany each submission. This would especially help newer grad students learn how to begin to do this sort of research. Maybe doing…

The problem is that reproducing something is really, really hard! Even if something doesn't reproduce in one experiment, it might be due to slight changes in some variables we don't even think about. There are some ways to circumvent it (e.g. team that's being reproduced cooperating with reproducing team and agreeing on what variables are important for the experiemnt and which are not), but it's really hard. The solu…

Every time some easy "Reproducibility is hard / not worth the effort" I hear "The original research wasn't meaningful or valuable".

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#296
post #242

Earlier quoted context omitted.

I’m not sure that’s fair in this context. In the past, a single paper with questionable or falsified results at a top tier conference was big news. Something that casts doubt on the validity of 53 papers at a top AI conference is at least notable. > whose actual findings remain valid Remain valid according to who? The same group that missed hundreds of hallucinated citations?

Which of these papers had falsified results and not bad citations? What is the base rate of bad citations pre-AI? And finally yes. Peer review does not mean clicking every link in the footnotes to make sure the original paper didn't mislink, though I'm sure after this bruhaha this too will be automated.

> Peer review does not mean clicking every link in the footnotes

It wasn't just broken links, but citing authors like "lastname, firstname" and made up titles.

I have done peer reviews for a (non-AI) CS conference and did at least skim the citations. For papers related to my domain, I was familiar with most of the citations already, and looked into any that looked odd.

Being familiar with the state of the art is, in theory, what qualifies you to do peer reviews.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#297

Earlier quoted context omitted.

Grad students don’t get to publish a thesis on reproduction. Everyone from the undergraduate research assistant to the tenured professor with research chairs are hyper focused on “publishing” as much “positive result” on “novel” work as possible

Publishing a replication could be a prerequisite to getting the degree The question is, how can universities coordinate to add this requirement and gain status from it

Prerequisite required by who, and why is that entity motivated to design such a requirement? Universities also want more novel breakthrough papers to boast about and to outshine other universities in the rankings. And if one is honest, other researchers also get more excited about new ideas than a failed replication that may for a thousand different reasons and the original authors will argue you did something wrong, or evaluated in an unfair way, and generally publicly accusing other researchers of doing bad work won't help your career much. It's a small world, you'd be making enemies with people who will sit on your funding evaluation committees, hiring committees and it just generally leads to drama. Also papers are superseded so fast that people don't even care that a no longer state of the art paper may have been wrong. There are 5 newer ones that perform better and nobody uses the old one. I'm just stating how things actually are, I don't say that this is good, but when you say something "should" happen, think about who exactly is motivated to drive such a change.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#298

Earlier quoted context omitted.

Grad students don’t get to publish a thesis on reproduction. Everyone from the undergraduate research assistant to the tenured professor with research chairs are hyper focused on “publishing” as much “positive result” on “novel” work as possible

But that seems almost trivially solved. In software it's common to value independent verification - e.g. code review. Someone who is only focused on writing new code instead of careful testing, refactoring, or peer review is widely viewed as a shitty developer by their peers. Of course there's management to consider and that's where incentives are skewed, but we're talking about a different structure. Why wouldn't th…

Universities are not really motivated to slow down the research careers of their employees, on the contrary. They are very much interested in their employees making novel, highly cited publications and bringing in grants that those publications can lead to.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#299
post #238

I spot-checked one of the flagged papers (from Google, co-authored by a colleague of mine) The paper was https://openreview.net/forum?id=0ZnXGzLcOg and the problem flagged was "Two authors are omitted and one (Kyle Richardson) is added. This paper was published at ICLR 2024." I.e., for one cited paper, the author list was off and the venue was wrong. And this citation was mentioned in the background section of the pa…

The thing is, when you copy paste a bibliography entry from the publisher or from Google Scholar, the authors won't be wrong. In this case, it is. If I were to write a paper with AI, I would at least manage the bibliography by hand, conscious of hallucinations. The fact that the hallucination is in the bibliography is a pretty strong indicator that the paper was written entirely with AI.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#300
post #238

I spot-checked one of the flagged papers (from Google, co-authored by a colleague of mine) The paper was https://openreview.net/forum?id=0ZnXGzLcOg and the problem flagged was "Two authors are omitted and one (Kyle Richardson) is added. This paper was published at ICLR 2024." I.e., for one cited paper, the author list was off and the venue was wrong. And this citation was mentioned in the background section of the pa…

Great job! I've tried to test their tool as well, but was totally paywalled.
Post reply on HN