Live data from Hacker News

GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

gptzero.me

261–270 of 528 posts

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#261

Earlier quoted context omitted.

In theory, asking grad students and early career folks to run replications would be a great training tool. But the problem isn’t just funding, it’s time. Successfully running a replication doesn’t get you a publication to help your career.

Yeah, but doesn't publishing an easily falsifiable paper end one?

The vast majority of papers is so insignifcant, nobody bothers to try and use and thereby replicate it.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#262
The authors talk about "a model's ability to align with human decisions" as a matter of the past. The omission in the paper is RLHF (Reinforcement Learning from Human Feedback). All these companies are "teaching machines to predict the preferences of people who click 'Accept All Cookies' without reading," by using low-paid human evaluators — “AI teachers.”

If we go back to Google, before its transformation into an AI powerhouse — as it gutted its own SERPs, shoving traditional blue links below AI-generated overlords that synthesize answers from the web’s underbelly, often leaving publishers starving for clicks in a zero-click apocalypse — what was happening?

The same kind of human “evaluators” were ranking pages. Pushing garbage forward. The same thing is happening with AI. As much as the human "evaluators" trained search engines to elevate clickbait, the very same humans now train large language models to mimic the judgment of those very same evaluators. A feedback loop of mediocrity — supervised by the... well, not the best among us. The machines still, as Stephen Wolfram wrote, for any given sequence, use the same probability method (e.g., “The cat sat on the...”), in which the model doesn’t just pick one word. It calculates a probability score for every single word in its vast vocabulary (e.g., “mat” = 40% chance, “floor” = 15%, “car” = 0.01%), and voilà! — you have a “creative” text: one of a gazillion mindlessly produced, soulless, garbage “vile bile” sludge emissions that pollute our collective brains and render us a bunch of idiots, ready to swallow any corporate poison sent our way.

In my opinion, even worse: the corporates are pushing toward “safety” (likely from lawsuits), and the AI systems are trained to sell, soothe, and please — not to think, or enhance our collective experience.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#264
So the headline says

>GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

And I'm left wondering if they mean 100 papers or 100 hallucinations

The subheading says

>GPTZero's analysis 4841 papers accepted by NeurIPS 2025 show there are at least 100 with confirmed hallucinations

Which accidentally a word, but seems to clarify that they do legitimately mean 100 papers.

A later heading says

>Table of 100 Hallucinated Citations in Published Across 53 NeurIPS Papers

Which suggests either the opposite, or that they chose a subset of their findings to point out a coincidentally similar number of incidents.

How many papers did they find hallucinations in? I'm still not certain. Is it 100, 53 or some other number altogether? Does their quality of scrutiny match the quality of their communication. If they did in-fact find 100 Hallucinations in 53 papers, would the inconsistency against their claim of "papers accepted by NeurIPS 2025 show there are at least 100 with confirmed hallucinations" meet their own bar for a hallucination?

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#265
post #253

> we discovered 100s of hallucinated citations missed by the 3+ reviewers who evaluated each paper. This says just as much about the humans involved.

Well for one, it's definitely not the responsibility of the reviewers to check that all the citations exist. That would be insane.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#266

Earlier quoted context omitted.

Yeah, but doesn't publishing an easily falsifiable paper end one?

One, it doesnt damage your reputation as much as one would think. But two, and more importantly, no one is checking. Tree falls in the forest, no one hears, yadi-yada.

Here's a work from last year which was plagiarized. The rare thing about this work is it was submitted to ICLR, which opened reviews for both rejected and accepted works.

You'll notice you can click on author names and you'll get links to their various scholar pages but notably DBLP, which makes it easy to see how frequently authors publish with other specific authors.

Some of those authors have very high citation counts... in the thousands, with 3 having over 5k each (one with over 18k).

https://openreview.net/forum?id=cIKQp84vqN

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#267
post #241

Earlier quoted context omitted.

If we are aiming for quality, then being wrong absolutely should be. I would argue that is how it works in real life anyway. What we quibble over is what is the appropriate cutoff.

There's a big gulf between being wrong because you or a collaborator missed an uncontrolled confounding factor and falsifying or altering results. Science accepts that people sometimes make mistakes in their work because a) they can also be expected to miss something eventually and b) a lot of work is done by people in training in labs you're not directly in control of (collaborators). They already aim for quality an…

In fields like psychology, though, you can be wrong for decades. If your result is foundational enough, and other people have "replicated" it, then most researchers will toss out contradictory evidence as "guess those people were an unrepresentative sample". This can be extremely harmful when, for instance, the prevailing view is "this demographic are just perverts" or "most humans are selfish thieves at heart, held back by perceived social consensus" – both examples where researcher misconduct elevated baseless speculation to the position of "prevailing understanding", which led to bad policy, which had devastating impacts on people's lives.

(People are better about this in psychology, now: schoolchildren are taught about some of the more egregious cases, even before university, and individual researchers are much more willing to take a sceptical view of certain suspect classes of "prevailing understanding". The fact that even I, a non-psychologist, know about this, is good news. But what of the fields whose practitioners don't know they have this problem?)

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#268
post #238

I spot-checked one of the flagged papers (from Google, co-authored by a colleague of mine) The paper was https://openreview.net/forum?id=0ZnXGzLcOg and the problem flagged was "Two authors are omitted and one (Kyle Richardson) is added. This paper was published at ICLR 2024." I.e., for one cited paper, the author list was off and the venue was wrong. And this citation was mentioned in the background section of the pa…

>this error does make me pause to wonder how much of the rest of the paper used AI assistance

And this is what's operative here. The error spotted, the entire class of error spotted, is easily checked/verified by a non-domain expert. These are the errors we can confirm readily, with obvious and unmistakable signature of hallucination.

If these are the only errors, we are not troubled. However: we do not know if these are the only errors, they are merely a signature that the paper was submitted without being thoroughly checked for hallucinations. They are a signature that some LLM was used to generate parts of the paper and the responsible authors used this LLM without care.

Checking the rest of the paper requires domain expertise, perhaps requires an attempt at reproducing the authors' results. That the rest of the paper is now in doubt, and that this problem is so widespread, threatens the validity of the fundamental activity these papers represent: research.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#269
post #241

Earlier quoted context omitted.

There's a big gulf between being wrong because you or a collaborator missed an uncontrolled confounding factor and falsifying or altering results. Science accepts that people sometimes make mistakes in their work because a) they can also be expected to miss something eventually and b) a lot of work is done by people in training in labs you're not directly in control of (collaborators). They already aim for quality an…

In fields like psychology, though, you can be wrong for decades . If your result is foundational enough, and other people have "replicated" it, then most researchers will toss out contradictory evidence as "guess those people were an unrepresentative sample". This can be extremely harmful when, for instance, the prevailing view is "this demographic are just perverts" or "most humans are selfish thieves at heart, held…

Yeah like I said the soft validation by subsequent papers is more true in more baseline physical sciences because it involves fewer uncontrollable variables. That's why I mentioned 'hard' sciences in my post, messy humans are messy and make science waaay harder.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#270
post #195
post #150

Earlier quoted context omitted.

I'd personally like to see top conferences grow a "reproducibility" track. Each submission would be a short tech report that chooses some other paper to re-implement. Cap 'em at three pages, have a lightweight review process. Maybe there could be artifacts (git repositories, etc) that accompany each submission. This would especially help newer grad students learn how to begin to do this sort of research. Maybe doing…

The problem is that reproducing something is really, really hard! Even if something doesn't reproduce in one experiment, it might be due to slight changes in some variables we don't even think about. There are some ways to circumvent it (e.g. team that's being reproduced cooperating with reproducing team and agreeing on what variables are important for the experiemnt and which are not), but it's really hard. The solu…

That's fine! The tech report should talk about what the researchers tried and what didn't work. I think submissions to the reproducibility track shouldn't necessarily have to be positive to be accepted, and conversely, I don't think the presence of a negative reproduction should necessarily impact an author's career negatively.
Post reply on HN