Live data from Hacker News

GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

gptzero.me

361–370 of 528 posts

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#361
post #339

Earlier quoted context omitted.

> If these are the only errors, we are not troubled. However: we do not know if these are the only errors, they are merely a signature that the paper was submitted without being thoroughly checked for hallucinations. They are a signature that some LLM was used to generate parts of the paper and the responsible authors used this LLM without care. I am troubled by people using an LLM at all to write academic research p…

> It's a shoddy, irresponsible way to work. And also plagiarism, when you claim authorship of it. It reminds me of kids these days and their fancy calculators! Those new fangled doohickeys just aren't reliable, and the kids never realize that they won't always have a calculator on them! Everyone should just do it the good old fashioned way with slide rules! Or these darn kids and their unreliable sources like Wikiped…

One issue with this analogy is that calculators really are precise when used correctly. LLMs are not.

I do think they can be used in research but not without careful checking. In my own work I’ve found them most useful as search aids and brainstorming sounding boards.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#362
post #355

Earlier quoted context omitted.

> It's a shoddy, irresponsible way to work. And also plagiarism, when you claim authorship of it. It reminds me of kids these days and their fancy calculators! Those new fangled doohickeys just aren't reliable, and the kids never realize that they won't always have a calculator on them! Everyone should just do it the good old fashioned way with slide rules! Or these darn kids and their unreliable sources like Wikiped…

Im really not motivated by this argument; it seems a false equivalence. Its not merely a spell checker or removing some tedium. As a professional mathematician I used wikipedia all the time to lookup quick facts before verifying it myself or elsewhere. A calculator well; I can use an actual programming language. Up until this point neither of those tools were asvertised or used by people to entirely replace human inp…

I hate to sound like a 19 year old on Reddit but:

AI People: "AI is a completely unprecedented technology where its introduction is unlike the introduction of any other transformative technology in history! We must treat it totally differently!"

Also AI People: "You're worried about nothing, this is just like when people were worried about the internet."

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#363
post #353

Earlier quoted context omitted.

I was an area chair on the NeurIPS program committee in 1997. I just looked and it seems that we had 1280 submissions. At that time, we were ultimately capped by the book size that MIT Press was willing to put out - 150 8-page articles. Back in 1997 we were all pretty sure we were on to something big. I'm sure people made mistakes on their bibliographies at that time as well! And did we all really dig up and read Met…

I cited Watson and Crick '53 in my PhD thesis and I did go dig it up and read it. I had to go to the basement of the library, use some sort of weird rotating knob to move a heavy stack of journals over, find some large bound book of the year's journals, and navigate to the paper. When I got the page, it had been cut out by somebody previous and replaced with a photocopied verison . (I also invested a HUGE amount of m…

Totally! If you haven't burrowed in the stacks as a grad student, you missed out.

The real challenges there aren't the "biggies" above, though, it's the ones in obscure journals you have to get copies of by inter-library agreements. My PhD was in applied probability and I was always happy if there were enough equations so that I could parse out the French or Russian-language explanation nearby.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#364
post #238

I spot-checked one of the flagged papers (from Google, co-authored by a colleague of mine) The paper was https://openreview.net/forum?id=0ZnXGzLcOg and the problem flagged was "Two authors are omitted and one (Kyle Richardson) is added. This paper was published at ICLR 2024." I.e., for one cited paper, the author list was off and the venue was wrong. And this citation was mentioned in the background section of the pa…

Agreed.

What I find more interesting is how easy these errors are to introduce and how unlikely they are to be caught. As you point out, a DOI checker would immediately flag this. But citation verification isn’t a first-class part of the submission or review workflow today.

We’re still treating citations as narrative text rather than verifiable objects. That implicit trust model worked when volumes were lower, but it doesn’t seem to scale anymore

There’s a project I’m working on at Duke University, where we are building a system that tries to address exactly this gap by making references and review labor explicit and machine verifiable at the infrastructure level. There’s a short explainer here that lays out what we mean, if useful context helps: https://liberata.info/

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#365
post #2

Yuck, this is going to really harm scientific research. There is already a problem with papers falsifying data/samples/etc, LLMs being able to put out plausible papers is just going to make it worse. On the bright side, maybe this will get the scientific community and science journalists to finally take reproducibility more seriously. I'd love to see future reporting that instead of saying "Research finds amazing che…

Yeah, spot on. If all we do is add more plausible sounding text on top of already fragile review and incentive structures, that really could make things worse rather than better

Your second point is the important one. AI may be the thing that finally forces the community to take reproducibility, attribution, and verification seriously. That’s very much the motivation behind projects like Liberata, which try to shift publishing away from novelty first narratives and toward explicit credit for replication, verification, and followthrough. If that cultural shift happens, this moment might end up being a painful but necessary correction.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#366
post #339

Earlier quoted context omitted.

> If these are the only errors, we are not troubled. However: we do not know if these are the only errors, they are merely a signature that the paper was submitted without being thoroughly checked for hallucinations. They are a signature that some LLM was used to generate parts of the paper and the responsible authors used this LLM without care. I am troubled by people using an LLM at all to write academic research p…

> It's a shoddy, irresponsible way to work. And also plagiarism, when you claim authorship of it. It reminds me of kids these days and their fancy calculators! Those new fangled doohickeys just aren't reliable, and the kids never realize that they won't always have a calculator on them! Everyone should just do it the good old fashioned way with slide rules! Or these darn kids and their unreliable sources like Wikiped…

I doubt that it's common for anyone to read a research paper and then question whether the researcher's calculator was working reliably.

Sure, maybe someday LLMs will be able to report facts in a mostly reliable fashion (like a typical calculator), but we're definitely not even close to that yet, so until we are the skepticism is very much warranted. Especially when the details really do matter, as in scientific research.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#367
post #340

Earlier quoted context omitted.

Why don't you look at the actual article? There are several more egregious examples, e.g., the authors being cited as "John Smith and Jane Doe"

I can see that either way. It could also be a placeholder until the actual author list is inserted. This could happen if you know the title, but not the authors and insert a temporary reference entry.

The first Doe and Smith example I could give that to (the title is real and the arxiv ID they give is "arXiv:2401.00001", which is definitely placeholder), but the second one doesn't match a title and has fake URL/DOI that don't actually go anywhere. There's a few that are unambiguously placeholders, but they really should have been caught in review for a conference this high up.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#368
post #238

I spot-checked one of the flagged papers (from Google, co-authored by a colleague of mine) The paper was https://openreview.net/forum?id=0ZnXGzLcOg and the problem flagged was "Two authors are omitted and one (Kyle Richardson) is added. This paper was published at ICLR 2024." I.e., for one cited paper, the author list was off and the venue was wrong. And this citation was mentioned in the background section of the pa…

Agreed. What I find more interesting is how easy these errors are to introduce and how unlikely they are to be caught. As you point out, a DOI checker would immediately flag this. But citation verification isn’t a first-class part of the submission or review workflow today. We’re still treating citations as narrative text rather than verifiable objects. That implicit trust model worked when volumes were lower, but it…

Citation checks are a workflow problem, not a model problem. Treat every reference as a dependency that must resolve and be reproducible. If the checker cannot fetch and validate it, it does not ship.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#369
post #361

Earlier quoted context omitted.

> It's a shoddy, irresponsible way to work. And also plagiarism, when you claim authorship of it. It reminds me of kids these days and their fancy calculators! Those new fangled doohickeys just aren't reliable, and the kids never realize that they won't always have a calculator on them! Everyone should just do it the good old fashioned way with slide rules! Or these darn kids and their unreliable sources like Wikiped…

One issue with this analogy is that calculators really are precise when used correctly. LLMs are not. I do think they can be used in research but not without careful checking. In my own work I’ve found them most useful as search aids and brainstorming sounding boards.

One issue with this analogy is that paper encyclopedias really are precise when used correctly. Wikipedia is not.

I do think it can be used in research but not without careful checking. In my own work I've found it most useful as a search aid and for brainstorming.

^ this same comment 10 years ago

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#370
post #28

NeurIPS leadership doesn’t think hallucinated references are necessarily disqualifying; see the full article from Fortune for a statement from them: https://archive.ph/yizHN > When reached for comment, the NeurIPS board shared the following statement: “The usage of LLMs in papers at AI conferences is rapidly evolving, and NeurIPS is actively monitoring developments. In previous years, we piloted policies regarding th…

> the content of the papers themselves are not necessarily invalidated. For example, authors may have given an LLM a partial description of a citation and asked the LLM to produce bibtex (a formatted reference) Maybe I'm overreacting, but this feels like an insanely biased response. They found the one potentially innocuous reason and latched onto that as a way to hand-wave the entire problem away. Science already had…

I don’t read the NeurIPS statement as malicious per se, but I do think it’s incomplete

They’re right that a citation error doesn’t automatically invalidate the technical content of a paper, and that there are relatively benign ways these mistakes get introduced. But focusing on intent or severity sidesteps the fact that citations, claims, and provenance are still treated as narrative artifacts rather than things we systematically verify

Once that’s the case, the question isn’t whether any single paper is “invalid” but whether the workflow itself is robust under current incentives and tooling.

A student group at Duke has been trying to think about with Liberata, i.e. what publishing looks like if verification, attribution, and reproducibility are first class rather than best effort

They have a short explainer here that lays out the idea if useful context helps: https://liberata.info/

Post reply on HN