Live data from Hacker News

GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

gptzero.me

421–430 of 528 posts

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#421
post #28

NeurIPS leadership doesn’t think hallucinated references are necessarily disqualifying; see the full article from Fortune for a statement from them: https://archive.ph/yizHN > When reached for comment, the NeurIPS board shared the following statement: “The usage of LLMs in papers at AI conferences is rapidly evolving, and NeurIPS is actively monitoring developments. In previous years, we piloted policies regarding th…

>NeurIPS leadership doesn’t think hallucinated references are necessarily disqualifying

That seems ridiculous.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#422
post #401
post #339

Earlier quoted context omitted.

> If these are the only errors, we are not troubled. However: we do not know if these are the only errors, they are merely a signature that the paper was submitted without being thoroughly checked for hallucinations. They are a signature that some LLM was used to generate parts of the paper and the responsible authors used this LLM without care. I am troubled by people using an LLM at all to write academic research p…

>also plagiarism To me, this is a reminder of how much of a specific minority this forum is. Nobody I know in real life, personally or at work, has expressed this belief. I have literally only ever encountered this anti-AI extremism (extremism in the non-pejorative sense) in places like reddit and here. Clearly, the authors in NeurIPS don't agree that using an LLM to help write is "plagiarism", and I would trust thei…

> AI Overview

> Plagiarism is using someone else's words, ideas, or work as your own without proper credit, a serious breach of ethics leading to academic failure, job loss, or legal issues, and can range from copying text (direct) to paraphrasing without citation (mosaic), often detected by software and best avoided by meticulous citation, quoting, and paraphrasing to show original thought and attribution.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#423
post #213

The innumeracy is load-bearing for the entire media ecosystem. If readers could do basic proportional reasoning, half of health journalism and most tech panic coverage would collapse overnight. GPTZero of course knows this. "100 hallucinations across 53 papers at prestigious conference" hits different than "0.07% of citations had issues, compared to unknown baseline, in papers whose actual findings remain valid."

> "0.07% of citations had issues

Nope, you are getting this part wrong. On purpose or by accident? Because it's pretty clear if you read the article they are not counting all citations that simply had issues. See "Defining Hallucinated Citations".

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#424

Earlier quoted context omitted.

> Those new fangled doohickeys just aren't reliable Except they are (unlike a chatbot, a calculator is perfectly deterministic), and the unreliability of LLMs is one of their most, if not the most , widespread target of criticism. Low effort doesn't even begin to describe your comment.

> Except they are (unlike a chatbot, a calculator is perfectly deterministic) LLM's are supposed to be stochastic. That is not a bug, I can see why you find that disappointing but it's just the reality of the tool. However, as I mentioned elsewhere calculators also have bugs and those bugs make their way into scientific research all the time. Floating point errors are particularly common, as are order of operations p…

That’s a strange argument. There are plenty of stochastic processes that have perfectly acceptable guarantees. A good example is Karger’s min-cut algorithm. You might not know what you get on any given single run, but you know EXACTLY what you’re going to get when you crank up the number of trials.

Nobody can tell you what you are going to get when you run an LLM once. Nobody can tell you what you’re going to get when you run it N times. There are, in fact, no guarantees at all. Nobody even really knows why it can solve some problems and why it can’t solve other except maybe it memorized the answer at some point. But this is not how they are marketed.

They are marketed as wondrous inventions that can SOLVE EVERYTHING. This is obviously not true. You can verify it yourself, with a simple deterministic problem: generate an arithmetic expression of length N. As you increase N, the probability that an LLM can solve it drops to zero.

Ok, fine. This kind of problem is not a good fit for an LLM. But which is? And after you’ve found a problem that seems like a good fit, how do you know? Did you test it systematically? The big LLM vendors are fudging the numbers. They’re testing on the training set, they’re using ad hoc measurements and so on. But don’t take my word for it. There’s lots of great literature out there that probes the eccentricities of these models; for some reason this work rarely makes its way into the HN echo chamber.

Now I’m not saying these things are broken and useless. Far from it. I use them every day. But I don’t trust anything they produce, because there are no guarantees, and I have been burned many times. If you have not been burned, you’re either exceptionally lucky, you are asking it to solve homework assignments, or you are ignoring the pain.

Excel bugs are not the same thing. Most of those problems can be found trivially. You can find them because Excel is a language with clear rules (just not clear to those particular users). The problem with Excel is that people aren’t looking for bugs.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#425

This feels less like scientific integrity and more like predatory marketing. I find this public "shame list" approach by GPTZero deeply unethical and technically suspect for several reasons: 1. Doxxing disguised as specific criticism: Publishing the names of authors and papers without prior private notification or independent verification is not how academic corrections work. It looks like a marketing stunt to genera…

Don't expect ethics from GPTZero. If you upload a large document, they'll give a fake 100% AI rating behind a blur until you pay up to get the actual analysis. This clearly serves to prey on paranoid authors who are worried about being perceived as using AI.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#426
post #401
post #339

Earlier quoted context omitted.

> If these are the only errors, we are not troubled. However: we do not know if these are the only errors, they are merely a signature that the paper was submitted without being thoroughly checked for hallucinations. They are a signature that some LLM was used to generate parts of the paper and the responsible authors used this LLM without care. I am troubled by people using an LLM at all to write academic research p…

>also plagiarism To me, this is a reminder of how much of a specific minority this forum is. Nobody I know in real life, personally or at work, has expressed this belief. I have literally only ever encountered this anti-AI extremism (extremism in the non-pejorative sense) in places like reddit and here. Clearly, the authors in NeurIPS don't agree that using an LLM to help write is "plagiarism", and I would trust thei…

The LLM model and version should be included as an author so there's useful information about where the content came from.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#427
post #309
post #238

I spot-checked one of the flagged papers (from Google, co-authored by a colleague of mine) The paper was https://openreview.net/forum?id=0ZnXGzLcOg and the problem flagged was "Two authors are omitted and one (Kyle Richardson) is added. This paper was published at ICLR 2024." I.e., for one cited paper, the author list was off and the venue was wrong. And this citation was mentioned in the background section of the pa…

> So the citation was not fabricated, but it was incorrectly attributed (perhaps via use of an AI autocomplete). Well the title says ”hallucinations”, not ”fabrications”. What you describe sounds exactly like what AI builders call hallucinations.

Read the article. The author uses the word "fabricate" repeatedly to describe the situation where the wrong authors are in the citation.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#428
post #339

Earlier quoted context omitted.

> If these are the only errors, we are not troubled. However: we do not know if these are the only errors, they are merely a signature that the paper was submitted without being thoroughly checked for hallucinations. They are a signature that some LLM was used to generate parts of the paper and the responsible authors used this LLM without care. I am troubled by people using an LLM at all to write academic research p…

There are legitimate , non-cheating ways to use LLMs for writing. I often use the wrong verb forms ("They synthesizes the ..."), write "though" when it should be "although", and forget to comma-separate clauses. LLMs are perfect for that. Generating text from scratch, however, is wrong.

> I often ... write "though" when it should be "although"

That is a purely imaginary "error". Anywhere you can use 'although', you are free to use 'though' instead.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#429

Earlier quoted context omitted.

“Anti-AI extremism”? Seriously? Where does this bizarre impulse to dogmatically defend LLM output come from? I don’t understand it. If AI is a reliable and quality tool, that will become evident without the need to defend it - it’s got billions (trillions?) of dollars backstopping it. The skeptical pushback is WAY more important right now than the optimistic embrace.

The fact that there is absurd AI hype right now doesn't mean that we should let equally absurd bullshit pass on the other side of the spectrum. Having a reasonable and accurate discussion about the benefits, drawbacks, side effects, etc. is WAY more important right now than being flagrantly incorrect in either direction. Meanwhile this entire comment thread is about what appears to be, as fumi2026 points out in their…

“anti-ai sentiment”

No that’s a straw man, sorry. Skepticism is not the same thing as irrational rejection. It means that I don’t believe you until you’ve proven with evidence that what you’re saying is true.

The efficacy and reliability of LLMs requires proof. Ai companies are pouring extraordinary, unprecedented amounts of money into promoting the idea that their products are intelligent and trustworthy. That marketing push absolutely dwarfs the skeptical voices and that’s what makes those voices more important at the moment. If the researchers named have claims made against them that aren’t true, that should be a pretty easy thing for them to refute.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#430
post #238

I spot-checked one of the flagged papers (from Google, co-authored by a colleague of mine) The paper was https://openreview.net/forum?id=0ZnXGzLcOg and the problem flagged was "Two authors are omitted and one (Kyle Richardson) is added. This paper was published at ICLR 2024." I.e., for one cited paper, the author list was off and the venue was wrong. And this citation was mentioned in the background section of the pa…

Sorry, but blaming it on "AI autocomplete" is the dumbest excuse ever. Author lists come from BibTeX entries and while they often contains errors since they can come from many sources, they do not contain completely made up authors. I don't share your view that hallucinated citations are less damaging in background section. Background, related works, and introduction is the sections where citations most often show up…

I'm not blaming anything on anything, because I did not (nor did the authors) confirm the cause of any of these errors.

> I don't share your view that hallucinated citations are less damaging in background section.

Who exactly is damaged in this particular instance?

Post reply on HN