Live data from Hacker News

GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

gptzero.me

301–310 of 528 posts

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#301

Earlier quoted context omitted.

> I'd love to see future reporting that instead of saying "Research finds amazing chemical x which does y" you see "Researcher reproduces amazing results for chemical x which does y. First discovered by z". Most people (that I talk to, at least) in science agree that there's a reproducibility crisis. The challenge is there really isn't a good way to incentivize that work. Fundamentally (unless you're independent weal…

> The challenge is there really isn't a good way to incentivize that work. Ban publication of any research that hasn't been reproduced.

If we did that, CERN could not publish, because nobody else has the capabilities they do. Do we really want to punish CERN (which has a good track record of scientific integrity) because their work can't be reproduced? I think the model in many of these cases is that the lab publishing has to allow some number of postdocs or competitor labs to come to their lab and work on reproducing it in-house with the same reagents (biological experiments are remarkably fragile).

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#302

Earlier quoted context omitted.

In theory, asking grad students and early career folks to run replications would be a great training tool. But the problem isn’t just funding, it’s time. Successfully running a replication doesn’t get you a publication to help your career.

You may well know this, but I get the sense that it isn’t necessarily common knowledge, so I want to spell it out anyway: In a lot of cases, the salary for a grad student or tech is small potatoes next to the cost of the consumables they use in their work. For example,I work for a lab that does a lot of sequencing, and if we’re busy one tech can use 10k worth of reagents in a week.

We are on the comment section about an AI conference and up until the last few years material/hardware costs for computer science research was very cheap compare to other sciences like medicine, biology etc. where they use bespoke instruments and materials. In CS, up until very recently, all you needed was a good consumer PC for each grad student that lasted for many years. Nowadays GPU clusters are more needed but funding is generally not keeping up with that, so even good university labs are way underresourced on this front.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#303
post #291
post #279

Earlier quoted context omitted.

Let he who is without sin cast the first stone…

If there were real consequences, we wouldn't be forced to churn out buggy nonsense by our employers. So we'd be able to take the time to do the right thing. Bug free software is possible, the world just says its not worth it today.

>Bug free software is possible, ...

Mr. Turing and his halting problem would like to politely disagree with this assertion.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#305
post #238

I spot-checked one of the flagged papers (from Google, co-authored by a colleague of mine) The paper was https://openreview.net/forum?id=0ZnXGzLcOg and the problem flagged was "Two authors are omitted and one (Kyle Richardson) is added. This paper was published at ICLR 2024." I.e., for one cited paper, the author list was off and the venue was wrong. And this citation was mentioned in the background section of the pa…

I see your point, but I don’t see where the author makes any claims about the specifics of the hallucinations, or their impact on the papers’ broader validity. Indeed, I would have found the removal of supposed “innocuous” examples to be far more deceptive than simply calling a spade a spade, and allowing the data to speak for itself.

The author calls the mistakes "confirmed hallucinations" without proof (just more or less evidence). The data never "speak for itself." The author curates the data and crafts a story about it. This story presented here is very suggestive (even using the term "hallucination" is suggestive). But calling it "100 suspected hallucinations", or "25 very likely hallucinations" does less for the author's end goal: selling their service.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#306

Earlier quoted context omitted.

while i agree that "reproducibility is overrated", i went ahead and read your medium post. my feedback to you is, my summary of that writing: "mike_hearn's take on policy-adjacent writing conducted by public health officials and published in journals that interacted with mike_hearn's valid and common but nonetheless subjective political dispute about COVID-19." i don't know how any of that writing generalizes to othe…

Thanks for reading it, or scan reading it maybe. Of the 18 papers discussed in the essay here's what they're about in order: - Alzheimers - Cancer - Alzheimers - Skin lesions (first paper discussed in the linked blog post) - Epidemiology (COVID) - Epidemiology (COVID, foot and mouth disease, Zika) - Misinformation/bot studies - More misinformation/bot studies - Archaeology/history - PCR testing (in general, discussio…

it only takes one drop of talking about COVID to make it about politics haha

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#307
post #272

Earlier quoted context omitted.

> Bibtex are often also incorrectly generated ...and including the erroneous entry is squarely the author's fault. Papers should be carefully crafted, not churned out. I guess that makes me sweetly naive

That's not happening for a similar reason people do not bug-check every single line of every single third-party library in their code. It's a chore that costs valuable time that you can instead spend on getting the actual stuff done. What's really important is that the scientific contribution is 100% correct and solid. For the references, the "good enough" paradigm applies. They mustn't be complete bogus, like the re…

To be honest, validating bibliographies does not cost valuable time. Every research group will have their own bibtex file to which every paper the group ever cited is added.

Typically when you add it you get the info from another paper or copy the bibtex entry from Google scholar, but it's really at most 10 minutes work, more likely 2-5. Every paper might have 5-10 new entries in the bibliography, so that's 1 hour or less of work?

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#308
post #238

I spot-checked one of the flagged papers (from Google, co-authored by a colleague of mine) The paper was https://openreview.net/forum?id=0ZnXGzLcOg and the problem flagged was "Two authors are omitted and one (Kyle Richardson) is added. This paper was published at ICLR 2024." I.e., for one cited paper, the author list was off and the venue was wrong. And this citation was mentioned in the background section of the pa…

>this error does make me pause to wonder how much of the rest of the paper used AI assistance And this is what's operative here. The error spotted, the entire class of error spotted, is easily checked/verified by a non-domain expert. These are the errors we can confirm readily, with obvious and unmistakable signature of hallucination. If these are the only errors, we are not troubled. However: we do not know if these…

This seems like finding spelling errors and using them to cast the entire paper into doubt.

I am unconvinced that the particular error mentioned above is a hallucination, and even less convinced that it is a sign of some kind of rampant use of AI.

I hope to find better examples later in the comment section.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#309
post #238

I spot-checked one of the flagged papers (from Google, co-authored by a colleague of mine) The paper was https://openreview.net/forum?id=0ZnXGzLcOg and the problem flagged was "Two authors are omitted and one (Kyle Richardson) is added. This paper was published at ICLR 2024." I.e., for one cited paper, the author list was off and the venue was wrong. And this citation was mentioned in the background section of the pa…

> So the citation was not fabricated, but it was incorrectly attributed (perhaps via use of an AI autocomplete).

Well the title says ”hallucinations”, not ”fabrications”. What you describe sounds exactly like what AI builders call hallucinations.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#310
post #293

Earlier quoted context omitted.

I see your point, but I don’t see where the author makes any claims about the specifics of the hallucinations, or their impact on the papers’ broader validity. Indeed, I would have found the removal of supposed “innocuous” examples to be far more deceptive than simply calling a spade a spade, and allowing the data to speak for itself.

The point is that they should focus on the meaningful errors, not the automiation of meaningless errors.

Why these are meaningless? How do I know now that the whole paper is not a slop?
Post reply on HN