Live data from Hacker News

Over fifty new hallucinations in ICLR 2026 submissions

gptzero.me

391–400 of 442 posts

Re: Over fifty new hallucinations in ICLR 2026 submissions

#391

Earlier quoted context omitted.

One incorrect way to think of it is "LLMs will sometimes hallucinate when asked to produce content, but will provide grounded insights when merely asked to review/rate existing content". A more productive (and secure) way to think of it is that all LLMs are "evil genies" or extremely smart, adversarial agents. If some PhD was getting paid large sums of money to introduce errors into your work, could they still mislea…

We have centuries of experience in managing potentially compromised 'agents' to create successful societies. Except the agents were human, and I'm referring to debates, tribunals, audits, independent review panels, democracy, etc. I'm not saying the LLM hallucination problem is solved, I'm just saying there's a wonderful myriad of ways to assemble pseudo-intelligent chatbots into systems where the trustworthiness of…

> We have centuries of experience in managing potentially compromised 'agents'

Not this kind though. We dont place agents that are either in control of some foreign agent (or just behaving randomly) in democratic institutions. And when we do, look at what happens. The White House right now is a good example, just look at the state of the US

Re: Over fifty new hallucinations in ICLR 2026 submissions

#392

Earlier quoted context omitted.

> it’s not “some people”, it’s practically everyone that doesn’t understand how these tools work, and even some people that do. Again, true for most things. A lot of people are terrible drivers, terrible judge of their own character, and terrible recreational drug users. Does that mean we need to remove all those things that can be misused? I much rather push back on shoddy work no matter what source. I don't care if…

>A lot of people are terrible drivers, terrible judge of their own character, and terrible recreational drug users. Does that mean we need to remove all those things that can be misused? Uhh, yes??? We have completely reshaped our cities so that cars can thrive in them at the expense of people. We have laws and exams and enforcement all to prevent cars from being driven by irresponsible people. And most drugs are lit…

People need to be responsible for things they put their name on. End of story. No AI company claims their models are perfect and don’t hallucinate. But paper authors should at least verify every single character their submit.

Re: Over fifty new hallucinations in ICLR 2026 submissions

#393

Earlier quoted context omitted.

> And yet, we’re not supposed to criticize the tool or its makers? Exactly, they're not forcing anyone to use these things, but sometimes others (their managers/bosses) forced them to. Yet it's their responsibility for choosing the right tool for the right problem, like any other professional. If a carpenter shows up to put a roof yet their hammer or nail-gun can't actually put in nails, who'd you blame; the tool, th…

> If a carpenter shows up to put a roof yet their hammer or nail-gun can't actually put in nails, who'd you blame; the tool, the toolmaker or the carpenter? I would be unhappy with the carpenter, yes. But if the toolmaker was constantly over-promising (lying?), lobbying with governments, pushing their tools into the hands of carpenters, never taking responsibility, then I would also criticize the toolmaker. It’s also…

That’s not what is happening with AI companies, and you damn well know it.

Re: Over fifty new hallucinations in ICLR 2026 submissions

#394

Earlier quoted context omitted.

>A lot of people are terrible drivers, terrible judge of their own character, and terrible recreational drug users. Does that mean we need to remove all those things that can be misused? Uhh, yes??? We have completely reshaped our cities so that cars can thrive in them at the expense of people. We have laws and exams and enforcement all to prevent cars from being driven by irresponsible people. And most drugs are lit…

People need to be responsible for things they put their name on. End of story. No AI company claims their models are perfect and don’t hallucinate. But paper authors should at least verify every single character their submit.

Yes, but they don’t. So clearly AI is a foot gun. What are doing about it?

Re: Over fifty new hallucinations in ICLR 2026 submissions

#395
post #78

Earlier quoted context omitted.

Scientists who use LLMs to write a paper are crappy scientists indeed. They need to be held accountable, even ostracised by the scientific community. But something is missing from the picture. Why is it that they came up with this idea in the first place? Who could have been peddling the impression (not an outright lie - they are very careful) about LLMs being these almost sentient systems with emergent intelligence,…

> Where is the god damn cure for cancer the LLMs were supposed to invent? Assuming that cure is meant as hyperbole, how about https://www.biorxiv.org/content/10.1101/2025.04.14.648850v3 ? AI models being used for bad purposes doesn't preclude them being used for good purposes.

...No, it was not meant as a hyperbole, as we were literally being told that these models will be able to do all of our work. I won't settle for the bullshit incremental wins here and there we see occassionally - I attribute those essentially to the old 'infinite number of monkeys typing on the infinite number of typewriters producing "Crime and Peace". No. that's not it - we were promised a god damn revolution, no less. Again, where is the cure for cancer and post-scarcity society ? Where is the AGI we were promised for the 2025? Let's hold the ghouls promising all that accountable for a change.

Re: Over fifty new hallucinations in ICLR 2026 submissions

#396

Earlier quoted context omitted.

Yeah WTF? Both authors and reviewers are hidden. Is this comment just an attempt to whip up racist fervor?

Don't understand why you're being downvoted, here.

Because the second sentence is inflammatory.

The side comment is right, it's about low versus high trust societies. Even if GP made a mistake on which names are relevant, they're not being racist about it.

Re: Over fifty new hallucinations in ICLR 2026 submissions

#397

If a carpenter builds a crappy shelf “because” his power tools are not calibrated correctly - that’s a crappy carpenter, not a crappy tool. If a scientist uses an LLM to write a paper with fabricated citations - that’s a crappy scientist. AI is not the problem, laziness and negligence is. There needs to be serious social consequences to this kind of thing, otherwise we are tacitly endorsing it.

I don't see much crappy power tool provider throwing billions in marketing and product placement to make them used everywhere.

Re: Over fifty new hallucinations in ICLR 2026 submissions

#398

Earlier quoted context omitted.

>A lot of people are terrible drivers, terrible judge of their own character, and terrible recreational drug users. Does that mean we need to remove all those things that can be misused? Uhh, yes??? We have completely reshaped our cities so that cars can thrive in them at the expense of people. We have laws and exams and enforcement all to prevent cars from being driven by irresponsible people. And most drugs are lit…

People need to be responsible for things they put their name on. End of story. No AI company claims their models are perfect and don’t hallucinate. But paper authors should at least verify every single character their submit.

>No AI company claims their models are perfect and don’t hallucinate

You can't have it both ways. Either AIs are worth billions BECAUSE they can run mostly unsupervised or they are not. This is exactly like the AI driving system in Autopilot, sold as autonomous but reality doesn't live up to it.

Re: Over fifty new hallucinations in ICLR 2026 submissions

#399

Earlier quoted context omitted.

Do you still feel the same way if the froiznok method is an ANOVA table of a linear regression, with a log-transformed outcome? Should I reference Fisher, Galton, Newton, the first person to log transform an outcome in a regression analysis, the first person to log transform the particular outcome used in your paper, the R developers, and Gauss and Markov for showing that under certain conditions OLS is the best line…

Yeah, there is an interesting question there (always has been). When do you stop citing the paper for a specific model? Just to take some examples, is BiCGStab famous enough now that we can stop citing van der Vorst? Is the AdS/CFT correspondence well known enough that we can stop citing Maldacena? Are transformers so ubiquitous that we don't have to cite "Attention is all you need" anymore? I would be closer to yes…

Yeah, I didn't want to be contrary just for the sake of it, the heuristics you mention seem like good ones, and if followed would probably already cut down on quite a few superfluous references in most papers.

Re: Over fifty new hallucinations in ICLR 2026 submissions

#400
post #292

Surely this is gross professional misconduct? If one of my postdocs did this they would be at risk of being fired. I would certainly never trust them again. If I let it get through, I should be at risk. As a reviewer, if I see the authors lie in this way why should I trust anything else in the paper? The only ethical move is to reject immediately. I acknowledge mistakes and so on are common but this is different leag…

What field are you in?

In many fields it's gross professional misconduct only in theory. This sort of thing is very common and there's never any consequence. LLM-generated citations specifically are a new problem but citations of documents that don't support the claim, contradict it, have nothing to do with it or were retracted years ago have been an issue for a long time.

Gwern wrote about this here:

https://gwern.net/leprechaun

"A major source of [false claim] transmission is the frequency with which researchers do not read the papers they cite: because they do not read them, they repeat misstatements or add their own errors, further transforming the leprechaun and adding another link in the chain to anyone seeking the original source. This can be quantified by checking statements against the original paper, and examining the spread of typos in citations: someone reading the original will fix a typo in the usual citation, or is unlikely to make the same typo, and so will not repeat it. Both methods indicate high rates of non-reading"

I first noticed this during COVID and did some blogging about it. In public health it is quite common to do things like present a number with a citation, and then the paper doesn't contain that number anywhere in it, or it does but the number was an arbitrary assumption pulled out of thin air rather than the empirical fact it was being presented as.

It was also very common for papers to open by saying something like, "Epidemiological models are a powerful tool for predicting the spread of disease" with eight different citations, and every single citation would be an unvalidated model - zero evidence that any of the cited models were actually good at prediction.

Bad citations are hardly the worst problem with these fields, but when you see how widespread it is and that nobody within the institutions cares it does lead to the reaction you're having where you just throw your hands up and declare whole fields to be writeoffs.

Post reply on HN