Live data from Hacker News

GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

gptzero.me

351–360 of 528 posts

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#351
post #339

Earlier quoted context omitted.

>this error does make me pause to wonder how much of the rest of the paper used AI assistance And this is what's operative here. The error spotted, the entire class of error spotted, is easily checked/verified by a non-domain expert. These are the errors we can confirm readily, with obvious and unmistakable signature of hallucination. If these are the only errors, we are not troubled. However: we do not know if these…

> If these are the only errors, we are not troubled. However: we do not know if these are the only errors, they are merely a signature that the paper was submitted without being thoroughly checked for hallucinations. They are a signature that some LLM was used to generate parts of the paper and the responsible authors used this LLM without care. I am troubled by people using an LLM at all to write academic research p…

> It's a shoddy, irresponsible way to work. And also plagiarism, when you claim authorship of it.

It reminds me of kids these days and their fancy calculators! Those new fangled doohickeys just aren't reliable, and the kids never realize that they won't always have a calculator on them! Everyone should just do it the good old fashioned way with slide rules!

Or these darn kids and their unreliable sources like Wikipedia! Everyone knows that you need a nice solid reliable source that's made out of dead trees and fact checked but up to 3 paid professionals!

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#352
This feels like a big nothingburger to me. Try an analysis on conference submissions (perhaps even published papers) from 1995 for comparison, and one from 2005, one from 2015. I recall the typos/errors/ommissions because I reviewed for them and I used them. Even then: so what? If I could find the reference relatively easily and with enough confidence I was fine. Rarely I couldnt find it and contacted the author. The job of the reviewer (or even author) isnt to be a nitpicky editor—that’s the editor’s job. Editing does not happen until the final printed publication is near, and only for accepted papers, nowadays sometimes it never happens. Now that is a problem perhaps, but it has nothing to do with the authors’ use of LLMs.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#353
post #260

Earlier quoted context omitted.

> rather the tasks of detecting errors have become increasingly difficult because of the associated effects of GPT. Did the increase to submissions to NeurIPS from 2020 to 2025 happen because ChatGPT came out in November of 2022? Or was AI getting hotter and hotter during this period, thereby naturally increasing submissions to ... an AI conference?

I was an area chair on the NeurIPS program committee in 1997. I just looked and it seems that we had 1280 submissions. At that time, we were ultimately capped by the book size that MIT Press was willing to put out - 150 8-page articles. Back in 1997 we were all pretty sure we were on to something big. I'm sure people made mistakes on their bibliographies at that time as well! And did we all really dig up and read Met…

I cited Watson and Crick '53 in my PhD thesis and I did go dig it up and read it.

I had to go to the basement of the library, use some sort of weird rotating knob to move a heavy stack of journals over, find some large bound book of the year's journals, and navigate to the paper. When I got the page, it had been cut out by somebody previous and replaced with a photocopied verison.

(I also invested a HUGE amount of my time into my bibliography in every paper I've written as first author, curating a database and writing scripts to format in the various journal formats. This involved multiple independent checks from several sources, repeated several times.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#354
post #339

Earlier quoted context omitted.

> If these are the only errors, we are not troubled. However: we do not know if these are the only errors, they are merely a signature that the paper was submitted without being thoroughly checked for hallucinations. They are a signature that some LLM was used to generate parts of the paper and the responsible authors used this LLM without care. I am troubled by people using an LLM at all to write academic research p…

> It's a shoddy, irresponsible way to work. And also plagiarism, when you claim authorship of it. It reminds me of kids these days and their fancy calculators! Those new fangled doohickeys just aren't reliable, and the kids never realize that they won't always have a calculator on them! Everyone should just do it the good old fashioned way with slide rules! Or these darn kids and their unreliable sources like Wikiped…

Annoying dismissal.

In an academic paper, you condense a lot of thinking and work, into a writeup.

Why would you blow off the writeup part, and impose AI slop upon the reviewers and the research community?

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#355
post #339

Earlier quoted context omitted.

> If these are the only errors, we are not troubled. However: we do not know if these are the only errors, they are merely a signature that the paper was submitted without being thoroughly checked for hallucinations. They are a signature that some LLM was used to generate parts of the paper and the responsible authors used this LLM without care. I am troubled by people using an LLM at all to write academic research p…

> It's a shoddy, irresponsible way to work. And also plagiarism, when you claim authorship of it. It reminds me of kids these days and their fancy calculators! Those new fangled doohickeys just aren't reliable, and the kids never realize that they won't always have a calculator on them! Everyone should just do it the good old fashioned way with slide rules! Or these darn kids and their unreliable sources like Wikiped…

Im really not motivated by this argument; it seems a false equivalence. Its not merely a spell checker or removing some tedium.

As a professional mathematician I used wikipedia all the time to lookup quick facts before verifying it myself or elsewhere. A calculator well; I can use an actual programming language.

Up until this point neither of those tools were asvertised or used by people to entirely replace human input.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#356
The prevalence of hallucinations in the system is another signs for change in the system. The citations should be treated less like narrative context and more like verifiable objects

Better detectors, like the article implies, won’t solve the problem, since AI will likely keep improving

It’s about the fact that our publishing workflows implicitly assume good faith manual verification, even as submission volume and AI assisted writing explode. That assumption just doesn’t hold anymore

A student initiative at Duke University has been working on what it might look like to address this at the publishing layer itself, by making references, review labor, and accountability explicit rather than implicit

There’s a short explainer video for their system: https://liberata.info/

It’s hard to argue that the current status quo will scale, so we need novel solutions like this.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#357
post #340

Earlier quoted context omitted.

This seems like finding spelling errors and using them to cast the entire paper into doubt. I am unconvinced that the particular error mentioned above is a hallucination, and even less convinced that it is a sign of some kind of rampant use of AI. I hope to find better examples later in the comment section.

Why don't you look at the actual article? There are several more egregious examples, e.g., the authors being cited as "John Smith and Jane Doe"

I can see that either way. It could also be a placeholder until the actual author list is inserted. This could happen if you know the title, but not the authors and insert a temporary reference entry.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#358
post #260

Earlier quoted context omitted.

> rather the tasks of detecting errors have become increasingly difficult because of the associated effects of GPT. Did the increase to submissions to NeurIPS from 2020 to 2025 happen because ChatGPT came out in November of 2022? Or was AI getting hotter and hotter during this period, thereby naturally increasing submissions to ... an AI conference?

I was an area chair on the NeurIPS program committee in 1997. I just looked and it seems that we had 1280 submissions. At that time, we were ultimately capped by the book size that MIT Press was willing to put out - 150 8-page articles. Back in 1997 we were all pretty sure we were on to something big. I'm sure people made mistakes on their bibliographies at that time as well! And did we all really dig up and read Met…

> And did we all really dig up and read Metropolis, Rosenbluth, Rosenbluth, Teller, and Teller (1953)?

If you didn't, you are lying. Full stop.

If you cite something, yes, I expect that you, at least, went back and read the original citation.

The whole damn point of a citation is to provide a link for the reader. If you didn't find it worth the minimal amount of time to go read, then why would your reader? And why did you inflict it on them?

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#359
post #305

Earlier quoted context omitted.

I see your point, but I don’t see where the author makes any claims about the specifics of the hallucinations, or their impact on the papers’ broader validity. Indeed, I would have found the removal of supposed “innocuous” examples to be far more deceptive than simply calling a spade a spade, and allowing the data to speak for itself.

The author calls the mistakes "confirmed hallucinations" without proof (just more or less evidence). The data never "speak for itself." The author curates the data and crafts a story about it. This story presented here is very suggestive (even using the term "hallucination" is suggestive). But calling it "100 suspected hallucinations", or "25 very likely hallucinations" does less for the author's end goal: selling th…

Obviously a post on a startup's blog will be more editorialized than an academic paper. Still, this seems like an important discussion to have.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#360
post #303
post #291

Earlier quoted context omitted.

If there were real consequences, we wouldn't be forced to churn out buggy nonsense by our employers. So we'd be able to take the time to do the right thing. Bug free software is possible, the world just says its not worth it today.

>Bug free software is possible, ... Mr. Turing and his halting problem would like to politely disagree with this assertion.

You misread the comment and DR Turing's paper.

Getting all possible software correct is impossible, clearly. Getting all the software you release is more possible because you can choose not to release the software that it is too hard to prove correct.

Not that the suggestion is practical or likely, but your assertion that it is impossible is incorrect.

Post reply on HN