Should be extremely easy for AI to successfully detect hallucinated references as they are semi-structured data with an easily verifiable ground truth.
GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
71–80 of 528 posts
Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
#72This suggests that nobody was screening this papers in the first place—so is it actually significant that people are using LLMs in a setting without meaningful oversight? These clearly aren't being peer-reviewed, so there's no natural check on LLM usage (which is different than what we see in work published in journals).
Academic venues don't have enough reviewers. This problem isn't new, and as publication volumes increase, it's getting sharply worse. Consider the unit economics. Suppose NeurIPS gets 20,000 papers in one year. Suppose each author should expect three good reviews, so area chairs assign five reviewers per paper. In total, 100,000 reviews need to be written. It's a lot of work, even before factoring emergency reviewers…
Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
#73NeurIPS leadership doesn’t think hallucinated references are necessarily disqualifying; see the full article from Fortune for a statement from them: https://archive.ph/yizHN > When reached for comment, the NeurIPS board shared the following statement: “The usage of LLMs in papers at AI conferences is rapidly evolving, and NeurIPS is actively monitoring developments. In previous years, we piloted policies regarding th…
Maybe I'm overreacting, but this feels like an insanely biased response. They found the one potentially innocuous reason and latched onto that as a way to hand-wave the entire problem away.
Science already had a reproducibility problem, and it now has a hallucination problem. Considering the massive influence the private sector has on the both the work and the institutions themselves, the future of open science is looking bleak.
Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
#74Earlier quoted context omitted.
Nope. I am still reviewing papers that propose solutions based on a technique X, conveniently ignoring research from two years ago that shows that X cannot be used on its own. Both the paper I reviewed and the research showing X cannot be used are in the same venue!
does it seem to be legitimate ignorance or maybe folks pushing ahead regardless of x being disproved?
There is also the reality that "one paper" or "one study" can be found contradicted almost anything, so if you just went with "some other paper/study debunks my premise" then you'd end up producing nothing. Plus many inside know that there's a lot of slop out there that gets published, so they can (sometimes reasonably IMHO) dismiss that "one paper" even when they do know about it.
It's (mostly) not fraud or malicious intent or ignorance, it's (mostly) humans existing in the system in which they must live.
Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
#75Earlier quoted context omitted.
Kinda gives the whole game away, doesn’t it? “It doesn’t actually matter if the citations are hallucinated.” In fairness, NeurIPS is just saying out loud what everyone already knows. Most citations in published science are useless junk: it’s either mutual back-scratching to juice h-index, or it’s the embedded and pointless practice of overcitation, like “Human beings need clean water to survive (Franz, 2002)”. Really…
There should be a way to drop any kind of circular citation ring from the indexes.
Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
#76Earlier quoted context omitted.
I think a _single_ instance of an LLM hallucination should be enough to retract the whole paper and ban further submissions.
Going through a retraction and blacklisting process is also a lot of work -- collecting evidence, giving authors a chance to respond and mediate discussion, etc. Labor is the bottleneck. There aren't enough academics who volunteer to help organize conferences. (If a reader of this comment is qualified to review papers and wants to step up to the plate and help do some work in this area, please email the program chair…
Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
#77NeurIPS leadership doesn’t think hallucinated references are necessarily disqualifying; see the full article from Fortune for a statement from them: https://archive.ph/yizHN > When reached for comment, the NeurIPS board shared the following statement: “The usage of LLMs in papers at AI conferences is rapidly evolving, and NeurIPS is actively monitoring developments. In previous years, we piloted policies regarding th…
I think a _single_ instance of an LLM hallucination should be enough to retract the whole paper and ban further submissions.
For example, authors may have given an LLM a partial description of a citation and asked the LLM to produce bibtex
This is equivalent to a typo. I’d like to know which “hallucinations” are completely made up, and which have a corresponding paper but contain some error in how it’s cited. The latter I don’t think matters.Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
#78I guess GPTZero has such a tool. I'm confused why it isn't used more widely by paper authors and reviewers
Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
#79Earlier quoted context omitted.
Academic venues don't have enough reviewers. This problem isn't new, and as publication volumes increase, it's getting sharply worse. Consider the unit economics. Suppose NeurIPS gets 20,000 papers in one year. Suppose each author should expect three good reviews, so area chairs assign five reviewers per paper. In total, 100,000 reviews need to be written. It's a lot of work, even before factoring emergency reviewers…
Seems like using tooling like this to identify papers with fake citations and auto-rejecting them before they ever get in front of a reviewer would kill two birds with one stone.
Another problem is that conferences move slowly and it's hard to adjust the publication workflow in such an invasive way. CVPR only recently moved from Microsoft's CMT to OpenReview to accept author submissions, for example.
There's a lot of opportunity for innovation in this space, but it's hard when everyone involved would need to agree to switch to a different workflow.
(Not shooting you down. It's just complicated because the people who would benefit are far away from the people who would need to do the work to support it...)
Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
#80Earlier quoted context omitted.
If I steal hundreds of thousands of dollars (salary, plus research grants and other funds) and produce fake output, what do you think is appropriate? To me, it's no different than stealing a car or tricking an old lady into handing over her fidelity account. You are stealing, and society says stealing is a criminal act.
We have a civil court system to handle stuff like this already.
EDIT - The threshold amount varies. Sometimes it's as low as a few hundred dollars. However, the point stands on its own, because there's no universe where the sum in question is in misdemeanor territory.