Live data from Hacker News

GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

gptzero.me

71–80 of 528 posts

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#71
This is mostly an ad for their product. But I bet you can get pretty good results with a Claude Code agent using a couple simple skills.

Should be extremely easy for AI to successfully detect hallucinated references as they are semi-structured data with an easily verifiable ground truth.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#72
post #51

This suggests that nobody was screening this papers in the first place—so is it actually significant that people are using LLMs in a setting without meaningful oversight? These clearly aren't being peer-reviewed, so there's no natural check on LLM usage (which is different than what we see in work published in journals).

Academic venues don't have enough reviewers. This problem isn't new, and as publication volumes increase, it's getting sharply worse. Consider the unit economics. Suppose NeurIPS gets 20,000 papers in one year. Suppose each author should expect three good reviews, so area chairs assign five reviewers per paper. In total, 100,000 reviews need to be written. It's a lot of work, even before factoring emergency reviewers…

Seems like using tooling like this to identify papers with fake citations and auto-rejecting them before they ever get in front of a reviewer would kill two birds with one stone.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#73
post #28

NeurIPS leadership doesn’t think hallucinated references are necessarily disqualifying; see the full article from Fortune for a statement from them: https://archive.ph/yizHN > When reached for comment, the NeurIPS board shared the following statement: “The usage of LLMs in papers at AI conferences is rapidly evolving, and NeurIPS is actively monitoring developments. In previous years, we piloted policies regarding th…

> the content of the papers themselves are not necessarily invalidated. For example, authors may have given an LLM a partial description of a citation and asked the LLM to produce bibtex (a formatted reference)

Maybe I'm overreacting, but this feels like an insanely biased response. They found the one potentially innocuous reason and latched onto that as a way to hand-wave the entire problem away.

Science already had a reproducibility problem, and it now has a hallucination problem. Considering the massive influence the private sector has on the both the work and the institutions themselves, the future of open science is looking bleak.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#74

Earlier quoted context omitted.

Nope. I am still reviewing papers that propose solutions based on a technique X, conveniently ignoring research from two years ago that shows that X cannot be used on its own. Both the paper I reviewed and the research showing X cannot be used are in the same venue!

does it seem to be legitimate ignorance or maybe folks pushing ahead regardless of x being disproved?

IMHO, It's mostly ignorance coming a push/drive to "publish or perish." When the stakes are so high and output is so valued, and when reproducability isn't required, it disincentivizes thorough work. The system is set up in a way that is making it fail.

There is also the reality that "one paper" or "one study" can be found contradicted almost anything, so if you just went with "some other paper/study debunks my premise" then you'd end up producing nothing. Plus many inside know that there's a lot of slop out there that gets published, so they can (sometimes reasonably IMHO) dismiss that "one paper" even when they do know about it.

It's (mostly) not fraud or malicious intent or ignorance, it's (mostly) humans existing in the system in which they must live.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#75

Earlier quoted context omitted.

Kinda gives the whole game away, doesn’t it? “It doesn’t actually matter if the citations are hallucinated.” In fairness, NeurIPS is just saying out loud what everyone already knows. Most citations in published science are useless junk: it’s either mutual back-scratching to juice h-index, or it’s the embedded and pointless practice of overcitation, like “Human beings need clean water to survive (Franz, 2002)”. Really…

There should be a way to drop any kind of circular citation ring from the indexes.

It's tough because some great citations are hard to find/procure still. I sometimes refer to papers that aren't on the Internet (eg. old wonderful books / journals).

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#76
post #58

Earlier quoted context omitted.

I think a _single_ instance of an LLM hallucination should be enough to retract the whole paper and ban further submissions.

Going through a retraction and blacklisting process is also a lot of work -- collecting evidence, giving authors a chance to respond and mediate discussion, etc. Labor is the bottleneck. There aren't enough academics who volunteer to help organize conferences. (If a reader of this comment is qualified to review papers and wants to step up to the plate and help do some work in this area, please email the program chair…

That's exactly why the inclusion of a hallucinated reference is actually a blessing. Instead going back and forth with the fraudster, just tell them to find the paper. If they can't, case closed. Massive amount of time and money saved.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#77
post #28

NeurIPS leadership doesn’t think hallucinated references are necessarily disqualifying; see the full article from Fortune for a statement from them: https://archive.ph/yizHN > When reached for comment, the NeurIPS board shared the following statement: “The usage of LLMs in papers at AI conferences is rapidly evolving, and NeurIPS is actively monitoring developments. In previous years, we piloted policies regarding th…

I think a _single_ instance of an LLM hallucination should be enough to retract the whole paper and ban further submissions.

   For example, authors may have given an LLM a partial description of a citation and asked the LLM to produce bibtex
This is equivalent to a typo. I’d like to know which “hallucinations” are completely made up, and which have a corresponding paper but contain some error in how it’s cited. The latter I don’t think matters.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#78
I don't understand: why aren't there automated tools to verify citations' existence? The data for a citation has a structured styling (APA, MLA, Chicago) and paper metadata is available via e.g. a web search, even if the paper contents are not

I guess GPTZero has such a tool. I'm confused why it isn't used more widely by paper authors and reviewers

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#79
post #72
post #51

Earlier quoted context omitted.

Academic venues don't have enough reviewers. This problem isn't new, and as publication volumes increase, it's getting sharply worse. Consider the unit economics. Suppose NeurIPS gets 20,000 papers in one year. Suppose each author should expect three good reviews, so area chairs assign five reviewers per paper. In total, 100,000 reviews need to be written. It's a lot of work, even before factoring emergency reviewers…

Seems like using tooling like this to identify papers with fake citations and auto-rejecting them before they ever get in front of a reviewer would kill two birds with one stone.

It's not always possible to distinguish between fake citations and citations that are simply hard to find (e.g. wonderful old books that aren't on the Internet).

Another problem is that conferences move slowly and it's hard to adjust the publication workflow in such an invasive way. CVPR only recently moved from Microsoft's CMT to OpenReview to accept author submissions, for example.

There's a lot of opportunity for innovation in this space, but it's hard when everyone involved would need to agree to switch to a different workflow.

(Not shooting you down. It's just complicated because the people who would benefit are far away from the people who would need to do the work to support it...)

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#80
post #41

Earlier quoted context omitted.

If I steal hundreds of thousands of dollars (salary, plus research grants and other funds) and produce fake output, what do you think is appropriate? To me, it's no different than stealing a car or tricking an old lady into handing over her fidelity account. You are stealing, and society says stealing is a criminal act.

We have a civil court system to handle stuff like this already.

Stealing more than a few thousand dollars is a felony, and felonies are handled in criminal court, not civil.

EDIT - The threshold amount varies. Sometimes it's as low as a few hundred dollars. However, the point stands on its own, because there's no universe where the sum in question is in misdemeanor territory.

Post reply on HN