Live data from Hacker News

GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

gptzero.me

491–500 of 528 posts

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#491
post #424

Earlier quoted context omitted.

> Except they are (unlike a chatbot, a calculator is perfectly deterministic) LLM's are supposed to be stochastic. That is not a bug, I can see why you find that disappointing but it's just the reality of the tool. However, as I mentioned elsewhere calculators also have bugs and those bugs make their way into scientific research all the time. Floating point errors are particularly common, as are order of operations p…

That’s a strange argument. There are plenty of stochastic processes that have perfectly acceptable guarantees. A good example is Karger’s min-cut algorithm. You might not know what you get on any given single run, but you know EXACTLY what you’re going to get when you crank up the number of trials. Nobody can tell you what you are going to get when you run an LLM once. Nobody can tell you what you’re going to get whe…

> But I don’t trust anything they produce, because there are no guarantees

> Did you test it systematically?

Yes! That is exactly the right way to use them. For example, when I'm vibe coding I don't ask it to write code. I ask it to write unit tests. THEN I verify that the test is actually testing for the right things with my own eyeballs. THEN I ask it to write code that passes the unit tests.

Same with even text formatting. Sometimes I ask it to write a pydantic script to validate text inputs of "x" format. Often writing the text to specify the format is itself a major undertaking. Then once the script is working I ask for the text, and tell it to use the script to validate it. After that I can know that I can expect deterministic results, though it often takes a few tries for it to pass the validator.

You CAN get deterministic results, you just have to adapt your expectations to match what the tool is capable of instead of expecting your hammer to magically be a great screwdriver.

I do agree that the SOLVE EVERYTHING crowd are severely misguided, but so are the SOLVE NOTHING crowd. It's a tool, just use it properly and all will be well.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#492
post #379

Earlier quoted context omitted.

I would argue that an LLM is a perfectly sensible tool for structure-preserving machine translation from another language to English. (Where by "another language", you could also also substitute "very poor/non-fluent English." Though IMHO that's a bit silly, even though it's possible; there's little sense in writing in a language you only half know, when you'd get a less-lossy result from just writing in your native…

Autotranslating technical texts is very hard. After the translation, you muct check that all the technical words were translated correctly, instead of a fancy synonym that does not make sense. (A friend has an old book translated a long time ago (by a human) from Russian to Spanish. Instead of " complex numbers ", the book calls them " complicated numbers ". :) )

The convenient thing in this case (verification of translation of academic papers from the speaker's native language to English) is that the authors of the paper likely already 1. can read English to some degree, and 2. are highly likely to be familiar specifically with the jargon terms of their field in both their own language and in English.

This is because, even in countries with a different primary spoken language, many academic subjects, especially at a graduate level (masters/PhD programs — i.e. when publishing starts to matter), are still taught at universities at least partly in English. The best textbooks are usually written in English (with acceptably-faithful translations of these texts being rarer than you'd think); all the seminal papers one might reference are likely to be in English; etc. For many programs, the ability to read English to some degree is a requirement for attendance.

And yet these same programs are also likely to provide lectures (and TA assistance) in the country's own native language, with the native-language versions of the jargon terms used. And any collaborative work is likely to also occur in the native language. So attendees of such programs end up exposed to both the native-language and English-language terms within their field.

This means that academics in these places often have very little trouble in verifying the fidelity of translation of the jargon in their papers. It's usually all the other stuff in the translation that they aren't sure is correct. But this can be cheaply verified by handing the paper to any fluently-multilingual non-academic and asking them to check the translation, with the instruction to just ignore the jargon terms because they were already verified.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#493

At least in one case the authors claimed to use ChatGPT to "generate the citations after giving it author-year in-text citations, titles, or their paraphrases." They pasted the hallucinations in without checking. They've since responded with corrections to real papers that in most cases are very similar to the hallucination, lending credibility to their claim.[1] Not great, but to be clear this is different from fabr…

I counted 15 hallucinated citations. The authors explanation is plausible, but it is still 15 citations to works they clearly have not read. Any university teaches you that citing sources you personally have not verified supports you claim(s) is fraudulent. Apologizing is not enough, they should retract the article.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#495
post #443

Earlier quoted context omitted.

> Why is someone behaving questionably the authority on whether that's OK? Because they are not. Using AI to help writing is something literally every company is pushing for.

How is that relevant? Companies care very little about plagiarism, at least in the ethical sense (they do care if they think it's a legal risk, but that has turned out to not be the case with AI, so far at least).

What do you mean how is that relevant? Its a vast majority opinion in society that using ai to help you write is fine. Calling it "plagiarism" is a tiny minority online opinion.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#496
post #286

Earlier quoted context omitted.

Just to clarify, you didn't actually look up the publications it was citing? For example, you just stayed in ChatGPT web and used the resources it provided there? Not ridiculing you of course, but am just curious. The last paper I wrote a couple months back I had GPT search out the publications for me, but I would always open a new tab and retrieve the actual publication.

I didn't because I wasn't really doing anything serious to my mind, I think? basically felt like watching an episode of pbs spacetime, I think the difference is it's more like playing a video game while thinking you're watching an episode of spacetime, if that makes sense? I don't use chatgpt for me real work that much, and I'm not a scientist, so it was for me just mucking around, it pushed me slightly over a line i…

Okay this makes more sense now and thanks for the explanation.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#497
They're not halluciations. Don't anthropomorphise this nonsene, call it what it is because this is not a new problem: this is garbage data, and that garbage data should have been caught. Having a submission pipeline that verifies sources even exist (not that they're citing the right thing) is one the bare minimum responsibilities of a paid journal.

This has almost nothing to do with AI, and everything to do with a journal not putting in the trivial effort (given how much it costs to get published by them) required to ensure subject integrity. Yeah AI is the new garbage generator, but this problem isn't new, citation verification's been part of review ever since citations became a thing.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#498

Earlier quoted context omitted.

> I often ... write "though" when it should be "although" That is a purely imaginary "error". Anywhere you can use 'although', you are free to use 'though' instead.

Yeah, but you cannot use although anywhere you can use though, though.

That's true, but the one-way substitutability still means there is no such thing as "writ[ing] 'though' when it should be 'although'".

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#499
post #495

Earlier quoted context omitted.

How is that relevant? Companies care very little about plagiarism, at least in the ethical sense (they do care if they think it's a legal risk, but that has turned out to not be the case with AI, so far at least).

What do you mean how is that relevant? Its a vast majority opinion in society that using ai to help you write is fine. Calling it "plagiarism" is a tiny minority online opinion.

First of all, the very fact that companies need to encourage it shows that it is not already a majority opinion in society, it is a majority opinion among company management, which is often extremely unethical.

Secondly, even if it is true that it is a majority opinion in society doesn't mean it's right. Society at large often misunderstands how technology works and what risks it brings and what are its inevitable downstream effects. It was a majority opinion in society for decades or centuries that smoking is neutral to your health - that doesn't mean they were right.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#500
post #238

I spot-checked one of the flagged papers (from Google, co-authored by a colleague of mine) The paper was https://openreview.net/forum?id=0ZnXGzLcOg and the problem flagged was "Two authors are omitted and one (Kyle Richardson) is added. This paper was published at ICLR 2024." I.e., for one cited paper, the author list was off and the venue was wrong. And this citation was mentioned in the background section of the pa…

This is par for the course for GPTZero, which also falsely claims they can detect AI generated text, a fundamentally impossible task to do accurately.

I'm not going to bat for GPTZero, but I think it's clearly possible to identify some AI-written prose. Scroll through LinkedIn or Twitter replies and there are clear giveaways in tone, phrasing and repeated structures (it's not just X it's Y).

Not to say that you could ever feasibly detect all AI-generated text, but if it's possible for people to develop a sense for the tropes of LLM content then there's no reason you couldn't detect it algorithmically.

Post reply on HN