Live data from Hacker News

Over fifty new hallucinations in ICLR 2026 submissions

gptzero.me

271–280 of 442 posts

Re: Over fifty new hallucinations in ICLR 2026 submissions

#271
post #147

Earlier quoted context omitted.

Code correctness should be checked automatically with the CI and testsuite. New tests should be added. This is exactly what makes sure these stupid errors don't bother the reviewer. Same for the code formatting and documentation.

What exactly is the analogy you’re suggesting, using LLMs to verify the citations?

not OP, but that wouldn't really be necessary.

One could submit their bibtex files and expect bibtex citations to be verifiable using a low level checker.

Worst case scenario if your bibtex citation was a variant of one in the checker database you'd be asked to correct it to match the canonical version.

However, as others here have stated, hallucinated "citations" are actually the lesser problem. Citing irrelevant papers based on a fly-by reference is a much harder problem; this was present even before LLMs, but this has now become far worse with LLMs.

Re: Over fifty new hallucinations in ICLR 2026 submissions

#272

Earlier quoted context omitted.

Are you sure this even works? My understanding is that hallucinations are a result of physics and the algorithms at play. The LLM always needs to guess what the next word will be. There is never a point where there is a word that is 100% likely to occur next. The LLM doesn't know what "reliable" sources are, or "real knowledge". Everything it has is user text, there is nothing it knows that isn't user text. It doesn'…

Telling the LLM not to hallucinate reminds me of, "why don't they build the whole plane out of the black box???" Most people are just lazy and eager to take shortcuts, and this time it's blessed or even mandated by their employer. The world is about to get very stupid.

"Do not hallucinate" - seems to "work" for Apple [1]

[1] https://arstechnica.com/gadgets/2024/08/do-not-hallucinate-t...

Re: Over fifty new hallucinations in ICLR 2026 submissions

#273

Earlier quoted context omitted.

Citations are a key part of the paper. If the paper isn’t supported by the citations, it’s not a good paper.

Have you ever followed citations before? In my experience, they don't support what is being citated, saying the opposite or not even related. It's probably only 60%-ish that actually cite something relevant.

Well yes, but just because that’s bad doesn’t mean this isn’t far worse.

Re: Over fifty new hallucinations in ICLR 2026 submissions

#274
post #81

One wonders why this has not been largely fully automated. If we track those citations anyway. Surely we have database of them and most of them are easily matched there. So only outliers need to be checked either as new latest papers or mistakes which should be close enough to something or real fakes. Maybe there just is no incentive for this type of activity.

We do have these things and they are often wrong. Loads of the examples given look better than things I’ve seen in real databases on this kind of thing and I worked in this area for a decade.

Re: Over fifty new hallucinations in ICLR 2026 submissions

#275
Once upon a time, in a more innocent age, someone made a parody (of an even older Evangelical propaganda comic [1]) that imputed an unexpected motivation to cultists who worship eldritch horrors: https://www.entrelineas.org/pdf/assets/who-will-be-eaten-fir...

It occurred to me that this interpretation is applicable here.

[1] https://en.wikipedia.org/wiki/Chick_tract

Re: Over fifty new hallucinations in ICLR 2026 submissions

#276

If a carpenter builds a crappy shelf “because” his power tools are not calibrated correctly - that’s a crappy carpenter, not a crappy tool. If a scientist uses an LLM to write a paper with fabricated citations - that’s a crappy scientist. AI is not the problem, laziness and negligence is. There needs to be serious social consequences to this kind of thing, otherwise we are tacitly endorsing it.

I'm an industrial electrician. A lot of poor electrical work is visible only to a fellow electrician, and sometimes only another industrial electrician. Bad technical work requires technical inspectors to criticize. Sometimes highly skilled ones.

an old boss of mine used to say there are no stupid electricians found alive, as they self select darwin award style

Re: Over fifty new hallucinations in ICLR 2026 submissions

#277

Earlier quoted context omitted.

you have completely missed the point of the analogy. breaking the analogy beyond the point where it is useful by introducing non-generalising specifics is not a useful argument. Otherwise I can counter your more specific non-generalising analogy by introducing little green aliens sabotaging your imaginary CI with the same ease and effect.

I disagree you could do that and claim to be reasonable. But I agree, because I'd rather discuss the pragmatics and not bicker over the semantics about an analogy. Introducing a token error, is different from plagiarism, no? Someone wrote code that can't compile, is different from someone "stealing" proprietary code from some company, and contributing it to some FOSS repo? In order to assume good faith, you also need…

Sure but the focus here is on the reviewer not the author.

The point is what is expected as reasonable review before one can "sign their name on it".

"Lazy" (or possibly malicious) authors will always have incentives to cut corners as long as no mechanisms exist to reject (or even penalise) the paper on submission automatically. Which would be the equivalent of a "compiler error" in the code analogy.

Effectively the point is, in the absence of such tools, the reviewer can only reasonably be expected to "look over the paper" for high-level issues; catching such low-level issues via manual checks by reviewers has massively diminishing returns for the extra effort involved.

So I don't think the conference shaming the reviewers here in the absence of providing such tooling is appropriate.

Re: Over fifty new hallucinations in ICLR 2026 submissions

#278

Every single person who did this should be censured by their own institutions. Do it more than once? Lose job. End of story.

Some of the examples listed are using the wrong paper title for a real paper (titles can change over time), missing authors (I’ve seen this before on Google Scholar bibitex), misstatements of venue (huh this working paper I added to my bibliography two years ago got published now nice to know), and similar mistakes. This just tells me you hate academics and want to hurt them gratuitously.

Re: Over fifty new hallucinations in ICLR 2026 submissions

#279
post #150

The legal system has a word to describe AI "slop" --- it is called "negligence". And as the remedy starts being applied (aka "liability"), the enthusiasm for AI will start to wane. I wouldn't be surprised if some businesses ban the use of AI --- starting with law firms.

The legal system has a word to describe software bugs --- it is called "negligence". And as the remedy starts being applied (aka "liability"), the enthusiasm for software will start to wane. What if anything do you think is wrong with my analogy? I doubt most people here support strict liability for bugs in code.

Very good analogy indeed. With one modification it makes perfect sense:

> And as the remedy starts being applied (aka "liability"), the enthusiasm for sloppy and poorly tested software will start to wane.

Many of us use AI to write code these days, but the burden is still on us to design and run all the tests.

Re: Over fifty new hallucinations in ICLR 2026 submissions

#280

Earlier quoted context omitted.

Yes, that's what it means to be a professional, you take responsibility for the quality of your work.

Well, then what does this say of LLM engineers at literally any AI company in existence if they are delivering AI that is unreliable then? Surely, they must take responsibility for the quality of their work and not blame it on something else.

I feel like what "unreliable" means, depends on well you understand LLMs. I use them in my professional work, and they're reliable in terms of I'm always getting tokens back from them, I don't think my local models have failed even once at doing just that. And this is the product that is being sold.

Some people take that to mean that responses from LLMs are (by human standards) "always correct" and "based on knowledge", while this is a misunderstanding about how LLMs work. They don't know "correct" nor do they have "knowledge", they have tokens, that come after tokens, and that's about it.

Post reply on HN