Live data from Hacker News

Over fifty new hallucinations in ICLR 2026 submissions

gptzero.me

241–250 of 442 posts

Re: Over fifty new hallucinations in ICLR 2026 submissions

#241
post #227
post #150

Earlier quoted context omitted.

The legal system has a word to describe software bugs --- it is called "negligence". And as the remedy starts being applied (aka "liability"), the enthusiasm for software will start to wane. What if anything do you think is wrong with my analogy? I doubt most people here support strict liability for bugs in code.

I don't even think GP knows what negligence is. Generally the law allows people to make mistakes, as long as a reasonable level of care is taken to avoid them (and also you can get away with carelessness if you don't owe any duty of care to the party). The law regarding what level of care is needed to verify genAI output is probably not very well defined, but it definitely isn't going to be strict liability. The emot…

I don’t get it, tech people clearly have the most to gain from AI like Claude Code.

Re: Over fifty new hallucinations in ICLR 2026 submissions

#242
post #65

Earlier quoted context omitted.

I think this is a bit unfair. The carpenters are (1) living in world where there’s an extreme focus on delivering as quicklyas possible, (2) being presented with a tool which is promised by prominent figures to be amazing, and (3) the tool is given at a low cost due to being subsidized. And yet, we’re not supposed to criticize the tool or its makers? Clearly there’s more problems in this world than «lazy carpenters»?

Yes, that's what it means to be a professional, you take responsibility for the quality of your work.

Well, then what does this say of LLM engineers at literally any AI company in existence if they are delivering AI that is unreliable then? Surely, they must take responsibility for the quality of their work and not blame it on something else.

Re: Over fifty new hallucinations in ICLR 2026 submissions

#243

Earlier quoted context omitted.

I'm an industrial electrician. A lot of poor electrical work is visible only to a fellow electrician, and sometimes only another industrial electrician. Bad technical work requires technical inspectors to criticize. Sometimes highly skilled ones.

I’d love to hear some examples of poor electrical work that you’ve come across that’s often missed or not seen.

A couple had just moved in a house and called me to replace the ceiling fan in the living room. I pulled the flush mount cover down to start unhooking the wire nuts and noticed RG58 (coax cable). Someone had used the center conductor as the hot wire! I ended up running 12/2 Romex from the switch. There was no way in hell I could have hooked it back up the way it was. This is just one example I've come across.

Re: Over fifty new hallucinations in ICLR 2026 submissions

#244
post #131

Earlier quoted context omitted.

>It's not easy to hallucinate papers out of whole cloth, but LLMs can easily and confidently do it, quote paragraphs that don't exist, and do it tirelessly and at a pace unmatched by humans. But no one is claiming these papers were hallucinated whole, so I don't see how that's relevant. This study -- notably to sell an "AI detector", which is largely a laughable snake-oil field -- looked purely at the accuracy of cit…

I've zero interest in the AI tool, I'm discussing the broader problem. The references were made up, and this is easier and faster to do with LLMs than with humans. Easier to do inadvertently, too. As I said, LLMs are a force multiplier for fraud and inadvertent errors. So it's a big deal.

I think we should see a chart as % of “fabricated” references from past 20 years. We should see a huge increase after 2020-2021. Anyone has this chart data?

Re: Over fifty new hallucinations in ICLR 2026 submissions

#245

Earlier quoted context omitted.

>LLMs can actually make up for their negative contributions. They could go through all the references of all papers and verify them, They will just hallucinate their existence. I have tried this before

I don’t see why this would be the case with proper tool calling and context management. If you tell a model with blank context ‘you are an extremely rigorous reviewer searching for fake citations in a possibly compromised text’ then it will find errors. It’s this weird situation where getting agents to act against other agents is more effective than trying to convince a working agent that it’s made a mistake. Perhaps…

If you truly think that you have an effective solution to hallucinations, you will become instantly rich because literally no one out there has an idea for an economically and technologically feasible solution to hallucinations

Re: Over fifty new hallucinations in ICLR 2026 submissions

#246

Earlier quoted context omitted.

They explain in the article what they consider a proper citation, an erroneous one and an hallucination, in the section "Defining Hallucitations". They also say than they have many false positives, mostly real papers who are not available online. Thad said, i am also very curious of the result than their tool, would give to papers from the 2010's and before.

If you look at their examples in the "Defining Hallucitations" section, I'd say those could be 100% human errors. Shortening authors' names, leaving out authors, misattributing authors, misspelling or misremembering the paper title (or having an old preprint-title, as titles do change) are all things that I would fully expect to happen to anyone in any field were things get ever got published. Modern tools have made…

I mean, if you’re able to take the citation, find the cited work, and definitively state ‘looks like they got the title wrong’ or ‘they attributed the paper to the wrong authors’, that doesn’t sound like what people usually mean when they say a ‘hallucinated’ citation. Work that is lazily or poorly cited but nonetheless attempts to cite real work is not the problem. Work which gives itself false authority by claiming to cite works that simply do not exist is the main concern surely?

Re: Over fifty new hallucinations in ICLR 2026 submissions

#247
post #110

Earlier quoted context omitted.

Not really true nowadays. Stuff in whitepapers needs to be verifiable which is kinda difficult with hallucinations. Whether the students directly used LLMs or just read content online that was produced with them and cited after just shows how difficult these things made gathering information that's verifiable.

> Stuff in whitepapers needs to be verifiable which is kinda difficult with hallucinations. That's... gibberish. Anything you can do to verify a paper, you can do to verify the same paper with all citations scrubbed. Whether the citations support the paper, or whether they exist at all, just doesn't have anything to do with what the paper says.

I dont think you know how whitepapers work then

Re: Over fifty new hallucinations in ICLR 2026 submissions

#248

If a carpenter builds a crappy shelf “because” his power tools are not calibrated correctly - that’s a crappy carpenter, not a crappy tool. If a scientist uses an LLM to write a paper with fabricated citations - that’s a crappy scientist. AI is not the problem, laziness and negligence is. There needs to be serious social consequences to this kind of thing, otherwise we are tacitly endorsing it.

Yeah seriously. Using an LLM to help find papers is fine. Then you read them. Then you use a tool like Zotero or manually add citations. I use Gemini Pro to identify useful papers that I might not yet have encountered before. But, even when asking to restrict itself to Pubmed resources, it's citations are wonky, citing three different version sources of the same paper (citations that don't say what they said they'd d…

The problem isn't whether they have more or less hallucinations. The problem is that they have them. And as long as they hallucinate, you have to deal with that. It doesn't really matter how you prompt, you can't prevent hallucinations from happening and without manual checking, eventually hallucinations will slip under the radar because the only difference between a real pattern and a hallucinated one is that one exists in the world and the other one doesn't. This is not something you can really counter with more LLMs either as it is a problem intrinsic to LLMs

Re: Over fifty new hallucinations in ICLR 2026 submissions

#249
post #69

Earlier quoted context omitted.

When academics are graded based on number of papers this is the result.

The problem isn't only papers it's that the world of academic computer science coalesced around conference submissions instead of journal submissions. This isn't new and was an issue 30 years ago when I was in grad school. It makes the work of conference organizes the little block holding up the entire system.

Makes me grateful I'm in an area of CS where the "big" conferences are like 500 attendees.

Re: Over fifty new hallucinations in ICLR 2026 submissions

#250

Earlier quoted context omitted.

I don’t see why this would be the case with proper tool calling and context management. If you tell a model with blank context ‘you are an extremely rigorous reviewer searching for fake citations in a possibly compromised text’ then it will find errors. It’s this weird situation where getting agents to act against other agents is more effective than trying to convince a working agent that it’s made a mistake. Perhaps…

If you truly think that you have an effective solution to hallucinations, you will become instantly rich because literally no one out there has an idea for an economically and technologically feasible solution to hallucinations

For references, as the OP said, I don't see why it isn't possible. It's something that exists and is accessible (even if paywalled) or doesn't exist. For reasoning hallucinations are different.
Post reply on HN