Live data from Hacker News

GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

gptzero.me

511–520 of 528 posts

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#511

Earlier quoted context omitted.

Indeed. My son has been accused by bullshit AI detection as having used AI, and it has devastated his work quality. After being "disciplined" for using AI (when he didn't), he now intentionally tries to "dumb down" his writing so that it doesn't sound so much like AI. The result is he writes much worse. What a shitty, shitty outcome. I've even found myself leaving typos and things in (even on sites like HN) because i…

Stop using em dashes, the fancy quotes that can’t be easily typed. Stop using overused words like certainly and delve. Stop using LLM template slop like “it’s not X, it’s Y”. Stop always doing lists of 3s. We know you didn’t use to use so many emojis or bolded text. Also, AI really fking hates the exclamation mark so that’s a great proof of humanity! Most people getting flagged are getting flagged because they actual…

That's all good advice, but it's not enough. He never uses em dashes or emojis in papers, and in the past when using exclamation marks he had teachers say, "don't use these in academic papers, they're not appropriate." Also mac OS loves to use the fancy quotes by default so when he's writing on a Mac, it's a pain in the ass to use regular quotes. It seems absurd to me that you'd have to jump through that hoop anyway just so it doesn't look like AI.

We've had teachers show us the screenshot output from their AI tool and it flags on things like "vocabulary word unusual for grade level." In my early 20s when I was dating my now-wife, she had a great vocabulary and I admired her for it, so I spent a lot of effort improving my vocabulary (well worth it by the way). When my son was born I intentionally used "big words" all the time with him (and explained what they meant when he didn't know) in the hopes that he would have a naturally large vocabulary when he got older. It worked very well. He routinely uses words even in conversation that even his teachers don't know. He writes even better than he speaks. But now being a statistical outlier is punishing him.

It flags plenty of other things like direct quotes (which he puts in quotation marks as he should) and includes it in the "score", so a quote heavy paper will sometimes show something like "65% produced by AI". He uses Google Docs so we can literally go through the whole history and see him writing the paper through time.

> Most people getting flagged are getting flagged because they actually used AI and couldn’t even be bothered to manually deslop it.

I'm sure that's true, but it doesn't excuse people using an automated tool that they don't understand and messing with other people's lives because of it. Just like when some cloud provider decides that your workload looks too much like crypto mining or something so AI auto-bans your account and shuts off your stuff.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#512
post #492

Earlier quoted context omitted.

Autotranslating technical texts is very hard. After the translation, you muct check that all the technical words were translated correctly, instead of a fancy synonym that does not make sense. (A friend has an old book translated a long time ago (by a human) from Russian to Spanish. Instead of " complex numbers ", the book calls them " complicated numbers ". :) )

The convenient thing in this case (verification of translation of academic papers from the speaker's native language to English) is that the authors of the paper likely already 1. can read English to some degree, and 2. are highly likely to be familiar specifically with the jargon terms of their field in both their own language and in English. This is because, even in countries with a different primary spoken languag…

> with the native-language versions of the jargon terms used

It depends on the country. Here in Argentina we use a lot of loaned words for technical terms, but I think in Spain they like to translate everything.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#513
post #500

Earlier quoted context omitted.

This is par for the course for GPTZero, which also falsely claims they can detect AI generated text, a fundamentally impossible task to do accurately.

I'm not going to bat for GPTZero, but I think it's clearly possible to identify some AI-written prose. Scroll through LinkedIn or Twitter replies and there are clear giveaways in tone, phrasing and repeated structures (it's not just X it's Y). Not to say that you could ever feasibly detect all AI-generated text, but if it's possible for people to develop a sense for the tropes of LLM content then there's no reason yo…

> there's no reason you couldn't detect it algorithmically

For any real world classifier there is a precision/recall tradeoff. Do you care more about false positives or false negatives? If you choose to truly minimize false positives you should simply always predict negative.

For your example “it’s not just X it’s Y” I agree it’s a red flag. But the origin of the pattern is from human text which the LLM picked up on. So some people did (and likely still do) use that construction.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#514

Earlier quoted context omitted.

Part of replication is skeptical review. That's also part of the scientific method. If we're not doing thorough review and replication, we're not really doing science. It's a lot of faith in people incentivized to do sloppy or dishonest work. Edit: I just read your article linked upthread. It was really good. I don't think we disagree except I say we need to attempt the steps of science wherever sensible and there's…

Yes, that's true. In theory, by the time it gets to the replication stage a paper has already been reviewed. In practice a replication is often the first time a paper is examined adversarially. There might be a useful form of hybrid here, like paying professional skeptics to review papers. The peer review concept academia works on is a very naive setup of the sort you'd expect given the prevailing ideology ("from eac…

"There might be a useful form of hybrid here, like paying professional skeptics to review papers."

This is how the scientific method is described. It's what much of the public thinks their money is paying for. So, I'm definitely for doing it for real or not calling it science.

Even the amount of review I saw you do on papers on your blog seems to exceed what much peer review is doing. So, how can we treat things as science if that aren't even meeting that standard, much less replication?

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#515

Earlier quoted context omitted.

> Nobody I know in real life, personally or at work, has expressed this belief. TBF, most people in real life don't even know how AI works to any degree, so using that as an argument that parent's opinion is extreme is kind of circular reasoning. > I have literally only ever encountered this anti-AI extremism (extremism in the non-pejorative sense) in places like reddit and here. I don't see parent's opinions as anti…

"If much of your research paper can be written by AI, I call into question whether or not it represents actual research" And what happens to this statement if next year or later this year the papers that can be autonomously written passes median human paper mark?

What does it mean to cross the median human paper mark? How os that measured?

It seems to me like most of the LLM benchmarks wind up being gamed. So, even if there were a good benchmark there, which I do not believe there is, the validity of the benchmark would likely diminish pretty quickly.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#516

Earlier quoted context omitted.

> Nobody I know in real life, personally or at work, has expressed this belief. TBF, most people in real life don't even know how AI works to any degree, so using that as an argument that parent's opinion is extreme is kind of circular reasoning. > I have literally only ever encountered this anti-AI extremism (extremism in the non-pejorative sense) in places like reddit and here. I don't see parent's opinions as anti…

> Research is supposed to be new ideas. If much of your research paper can be written by AI, I call into question whether or not it represents actual research. One would hope the authors are forming a hypothesis, performing an experiment, gathering and analysing results, and only then passing it to the AI to convert it into a paper. If I have a theory that, IDK, laser welds in a sine wave pattern are stronger than la…

I am not an academic, so correct me if I am wrong, but in your example, the actual writing would probably only represent a small fraction of the time spent. Is it even worth using AI for anything other than spelling and grammar correction at that point? I think using an LLM to generate a paper from high level points wouldn't save much, if any, time if it was then reviewed the way that would require.

My brother in law is a professor, and he has a pretty bad opinion of colleagues that use LLMs to write papers, as his field (economics) doesn't involve much experimentation, and instead relies on data analysis, simulation, and reasoning. It seemed to me like the LLM assisted papers that he's seen have mostly been pretty low impact filler papers.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#517

Earlier quoted context omitted.

That's not happening for a similar reason people do not bug-check every single line of every single third-party library in their code. It's a chore that costs valuable time that you can instead spend on getting the actual stuff done. What's really important is that the scientific contribution is 100% correct and solid. For the references, the "good enough" paradigm applies. They mustn't be complete bogus, like the re…

To be honest, validating bibliographies does not cost valuable time. Every research group will have their own bibtex file to which every paper the group ever cited is added. Typically when you add it you get the info from another paper or copy the bibtex entry from Google scholar, but it's really at most 10 minutes work, more likely 2-5. Every paper might have 5-10 new entries in the bibliography, so that's 1 hour or…

The complaint here is about this being an insufficient amount of effort because the bibtex entry from Google Scholar is wrong sometimes.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#518

Earlier quoted context omitted.

The fact that there is absurd AI hype right now doesn't mean that we should let equally absurd bullshit pass on the other side of the spectrum. Having a reasonable and accurate discussion about the benefits, drawbacks, side effects, etc. is WAY more important right now than being flagrantly incorrect in either direction. Meanwhile this entire comment thread is about what appears to be, as fumi2026 points out in their…

“anti-ai sentiment” No that’s a straw man, sorry. Skepticism is not the same thing as irrational rejection. It means that I don’t believe you until you’ve proven with evidence that what you’re saying is true. The efficacy and reliability of LLMs requires proof. Ai companies are pouring extraordinary, unprecedented amounts of money into promoting the idea that their products are intelligent and trustworthy. That marke…

There was a front page post just a couple of days ago where the article claimed LLMs have not improved in any way in over a year - an obviously absurd statement. A year before Opus 4.5, I couldn't get models to spit out a one shot Tampermonkey script to add chapter turns to my arrow keys. Now I can one small personal projects in claude code.

If you are saying that people are not making irrational and intellectually dishonest arguments about AI, I can't believe that we're reading the same articles and same comments.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#519

Earlier quoted context omitted.

The fact that there is absurd AI hype right now doesn't mean that we should let equally absurd bullshit pass on the other side of the spectrum. Having a reasonable and accurate discussion about the benefits, drawbacks, side effects, etc. is WAY more important right now than being flagrantly incorrect in either direction. Meanwhile this entire comment thread is about what appears to be, as fumi2026 points out in their…

Isn’t that the whole point of publishing? This happened plenty before AI too, and the claims are easily verified by checking the claimed hallucinations. Don’t publish things that aren’t verified and you won’t have a problem, same as before but perhaps now it’s easier to verify, which is a good thing. We see this problem in many areas, last week it was a criminal case where a made up law was referenced, luckily the ju…

> Isn’t that the whole point of publishing?

No, obviously not. You're confusing a marketing post by people with a product to sell with an actual review of the work by the relevant community, or even review by interested laypeople.

This is a marketing post where they provide no evidence that any of these are hallucinations beyond their own AI tool telling them so - and how do we know it isn't hallucinating? Are there hallucinations in there? Almost certainly. Would the authors deserve being called out by people reviewing their work? Sure.

But what people don't deserve is an unrelated VC funded tech company jumping in and claiming all of their errors are LLM hallucinations when they have no actual proof, painting them all a certain way so they can sell their product.

> Don’t publish things that aren’t verified and you won’t have a problem

If we were holding this company to the same standard, this blog wouldn't be posted either. They have not and can not verify their claims - they can't even say that their claims are based on their own investigations.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#520

Earlier quoted context omitted.

>this error does make me pause to wonder how much of the rest of the paper used AI assistance And this is what's operative here. The error spotted, the entire class of error spotted, is easily checked/verified by a non-domain expert. These are the errors we can confirm readily, with obvious and unmistakable signature of hallucination. If these are the only errors, we are not troubled. However: we do not know if these…

This seems like finding spelling errors and using them to cast the entire paper into doubt. I am unconvinced that the particular error mentioned above is a hallucination, and even less convinced that it is a sign of some kind of rampant use of AI. I hope to find better examples later in the comment section.

> This seems like finding spelling errors and using them to cast the entire paper into doubt.

Well, to be fair, I did encounter this from actual human peer reviewers before the whole LLM thing. People do that.

Post reply on HN