Live data from Hacker News

GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

gptzero.me

21–30 of 528 posts

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#21
I was getting completely AI-generated reviews for a WACV publication back in 2024. The area chairs are so overworked that authors don't have much recourse, which sucks but is also really hard to handle unless more volunteers step up to the bat to help organize the conference.

(If you're qualified to review papers, please email the program chair of your favorite conference and let them know -- they really need the help!)

As for my review, the review form has a textbox for a summary, a textbox for strengths, a textbox for weaknesses, and a textbox for overall thoughts. The review I received included one complete set of summary/strengths/weaknesses/closing thoughts in the summary text box, another distinct set of summary/strengths/weaknesses/closing thoughts in the strengths, another complete and distinct review in the weaknesses, and a fourth complete review in the closing thoughts. Each of these four reviews were slightly different and contradicted each other.

The reviewer put my paper down as a weak reject, but also said "the pros greatly outweigh the cons."

They listed "innovative use of synthetic data" as a strength, and "reliance on synthetic data" as a weakness.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#22

It is very concerning that these hallucinations passed through peer review. It's not like peer review is a fool-proof method or anything, but the fact that reviewers did not check all references and noticed clearly bogus ones is alarming and could be a sign that the article authors weren't the only ones using LLMs in the process...

Is it common for peer reviewers to check references? Somehow I thought they mostly focused on whether the experiment looked reasonable and the conclusions followed.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#23
post #10
post #3

It would be great if those scientists who use AI without disclosing it get fucked for life.

Harsh sentiment. Pretty soon every knowledge worker will use AI every day. Should people disclose spellcheckers powered by AI? Disclosing is not useful. Being careful in how you use it and checking work is what matters.

> Should people disclose spellcheckers powered by AI?

Thank you for that perfect example of a strawman argument! No, spellcheckers that use AI is not the main concern behind disclosing the use of AI in generating scientific papers, government reports, or any large block of nonfiction text that you paid for that is supposed to make to sense.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#24
post #2

Yuck, this is going to really harm scientific research. There is already a problem with papers falsifying data/samples/etc, LLMs being able to put out plausible papers is just going to make it worse. On the bright side, maybe this will get the scientific community and science journalists to finally take reproducibility more seriously. I'd love to see future reporting that instead of saying "Research finds amazing che…

Have they solved the issue where papers that cite research already invalidated are still being cited?

AFAIK, no, but I could see there being cause to push citations to also cite the validations. It'd be good if standard practice turned into something like

Paper A, by bob, bill, brad. Validated by Paper B by carol, clare, charlotte.

or

Paper A, by bob, bill, brad. Unvalidated.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#25
post #10
post #3

It would be great if those scientists who use AI without disclosing it get fucked for life.

Harsh sentiment. Pretty soon every knowledge worker will use AI every day. Should people disclose spellcheckers powered by AI? Disclosing is not useful. Being careful in how you use it and checking work is what matters.

People are accountable for the results they produce using AI. So a scientist is responsible for made up sources in their paper, which is plain fraud.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#26

This suggests that nobody was screening this papers in the first place—so is it actually significant that people are using LLMs in a setting without meaningful oversight? These clearly aren't being peer-reviewed, so there's no natural check on LLM usage (which is different than what we see in work published in journals).

As one who reviews 20+ papers per year, we don't have time to verify each reference.

We verify: is the stuff correct, and is it worthy of publication (in the given venue) given that it is correct.

There is still some trust in the authors to not submit made-up-stuff, albeit it is diminishing.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#28
NeurIPS leadership doesn’t think hallucinated references are necessarily disqualifying; see the full article from Fortune for a statement from them: https://archive.ph/yizHN

> When reached for comment, the NeurIPS board shared the following statement: “The usage of LLMs in papers at AI conferences is rapidly evolving, and NeurIPS is actively monitoring developments. In previous years, we piloted policies regarding the use of LLMs, and in 2025, reviewers were instructed to flag hallucinations. Regarding the findings of this specific work, we emphasize that significantly more effort is required to determine the implications. Even if 1.1% of the papers have one or more incorrect references due to the use of LLMs, the content of the papers themselves are not necessarily invalidated. For example, authors may have given an LLM a partial description of a citation and asked the LLM to produce bibtex (a formatted reference). As always, NeurIPS is committed to evolving the review and authorship process to best ensure scientific rigor and to identify ways that LLMs can be used to enhance author and reviewer capabilities.”

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#29

It is very concerning that these hallucinations passed through peer review. It's not like peer review is a fool-proof method or anything, but the fact that reviewers did not check all references and noticed clearly bogus ones is alarming and could be a sign that the article authors weren't the only ones using LLMs in the process...

Is it common for peer reviewers to check references? Somehow I thought they mostly focused on whether the experiment looked reasonable and the conclusions followed.

In journal publications it is, but without DOIs it's difficult.

In conference publications, it's less common.

Conference publications (like NEURips) is treated as announcement of results, not verified.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#30
post #15

Which is worse: a) p-hacking and suppressing null results b) hallucinations c) falsifying data Would be cool to see an analysis of this

All 3 of these should be categorized as fraud, and punished criminally.

criminally feels excessive?
Post reply on HN