Live data from Hacker News

Over fifty new hallucinations in ICLR 2026 submissions

gptzero.me

411–420 of 442 posts

Re: Over fifty new hallucinations in ICLR 2026 submissions

#411
post #78

Earlier quoted context omitted.

> Where is the god damn cure for cancer the LLMs were supposed to invent? Assuming that cure is meant as hyperbole, how about https://www.biorxiv.org/content/10.1101/2025.04.14.648850v3 ? AI models being used for bad purposes doesn't preclude them being used for good purposes.

...No, it was not meant as a hyperbole, as we were literally being told that these models will be able to do all of our work. I won't settle for the bullshit incremental wins here and there we see occassionally - I attribute those essentially to the old 'infinite number of monkeys typing on the infinite number of typewriters producing "Crime and Peace". No. that's not it - we were promised a god damn revolution, no l…

My understanding is that they're promising those as endgoals of the development trajectory, not that any current model actually is AGI. Did anyone really claim that, let's say GPT4, would cure cancer or meet any AGI standard?

Re: Over fifty new hallucinations in ICLR 2026 submissions

#412
post #334

It astonishes me that there would be so many cases of things like wrong authors. I began using a citation manager that extracted metadata automatically (zotero in my case) more than 15 years ago, and can’t imagine writing an academic paper without it or a similar tool. How are the authors even submitting citations? Surely they could be required to send a .bib or similar file? It’s so easy to then quality control at l…

Maybe you haven’t carefully checked yet the correctness of automatic tools or of the associated metadata. Zotero is certainly not bug free. Even authors themselves have miss-cited their own past work on occasion, and author lists have had errors that get revised upon resubmission or corrected in errata after publication. The DOI is indeed great, and if it is correct, I can still use the citation as a reader, but the…

I agree that the author lists in various metadata sources and databases are often a bit wrong (weird formatting of names for instance is very common), but many of the cases in the OP article are pretty egregious and far beyond just data entry issues.

Presumably the citation scanner they're using is relying on similar data sources as Zotero in any case to detect these sorts of issues.

Regardless, my comment still stands, it seems like the submission is relying on the actual text of the bibliography being correct, rather than requiring a machine readable citation metadata file of some sort, which would at least allow much of the quality control checks to be automated (and certainly would preclude complete hallucinations of nonexistent papers getting through).

Re: Over fifty new hallucinations in ICLR 2026 submissions

#413
post #241
post #227

Earlier quoted context omitted.

I don't even think GP knows what negligence is. Generally the law allows people to make mistakes, as long as a reasonable level of care is taken to avoid them (and also you can get away with carelessness if you don't owe any duty of care to the party). The law regarding what level of care is needed to verify genAI output is probably not very well defined, but it definitely isn't going to be strict liability. The emot…

I don’t get it, tech people clearly have the most to gain from AI like Claude Code.

Computer code is highly deterministic. This allows it to be tested fairly easily. Unfortunately, code productionn is not the only use-case for AI.

Most things in life are not as well defined --- a matter of judgment.

AI is being applied in lots of real world cases where judgment is required to interpret results. For example, "Does this patient have cancer". And it is fairly easy to show that AI's judgment can be highly suspect. There are often legal implications for poor judgment --- i.e. medical malpractice.

Maybe you can argue that this is a mis-application of AI --- and I don't necessarily disagree --- but the point is, once the legal system makes this abundantly clear, the practical business case for AI is going to be severely reduced if humans still have to vet the results in every case.

Re: Over fifty new hallucinations in ICLR 2026 submissions

#414

Earlier quoted context omitted.

Citations are a key part of the paper. If the paper isn’t supported by the citations, it’s not a good paper.

Have you ever followed citations before? In my experience, they don't support what is being citated, saying the opposite or not even related. It's probably only 60%-ish that actually cite something relevant.

I follow them a lot. I’ve also had cases where they don’t support the paper.

This doesn’t make it okay. Bad human writer and reviewer practices are also bad.

Re: Over fifty new hallucinations in ICLR 2026 submissions

#415

Earlier quoted context omitted.

...No, it was not meant as a hyperbole, as we were literally being told that these models will be able to do all of our work. I won't settle for the bullshit incremental wins here and there we see occassionally - I attribute those essentially to the old 'infinite number of monkeys typing on the infinite number of typewriters producing "Crime and Peace". No. that's not it - we were promised a god damn revolution, no l…

My understanding is that they're promising those as endgoals of the development trajectory, not that any current model actually is AGI. Did anyone really claim that, let's say GPT4, would cure cancer or meet any AGI standard?

Well, Sam Altman had said not long ago we would have AGI in 2025, and has been constantly implying something about "AI Scientists" and this and that. He literally said "We now know how to build AGI", also not long ago. He also stated that ChatGPT passed the Turing test without much fuss. The Anthropic has been pushing the narrative about the massive job loss, implying again that there would be an absolutely transforming impact coming soon. The Microsoft MBA-in-charge will have you believe his entire life and work is supposedly managed by an army of Clippy 2.0. The Google-MBA-in-charge has now started day-dreaming about space-based clusters, because guess what, his tool generates better fake pictures than Altman's. He too peddles the nonsense about the superpowerful AI. So, again yes, they said the AI would cure cancer and meet the AGI standard, so I demand they be held accountable for their own words and provide the answers to those questions!

Re: Over fifty new hallucinations in ICLR 2026 submissions

#416

If a carpenter builds a crappy shelf “because” his power tools are not calibrated correctly - that’s a crappy carpenter, not a crappy tool. If a scientist uses an LLM to write a paper with fabricated citations - that’s a crappy scientist. AI is not the problem, laziness and negligence is. There needs to be serious social consequences to this kind of thing, otherwise we are tacitly endorsing it.

"Anyone, from the most clueless amateur to the best cryptographer, can create an algorithm that he himself can’t break."--Bruce Schneier There's a corollary here with LLMs, but I'm not pithy enough to phrase it well. Anyone can create something using LLMs that they, themselves, aren't skilled enough to spot the LLMs' hallucinations. Or something. LLMs are incredibly good at exploiting peoples' confirmation biases. If…

> I do not believe there exists a way to safely use LLMs in scientific processes.

What about giving the LLM a narrowly scoped role as a hostile reviewer, while your job is to strengthen the write-up to address any valid objections it raises, plus any hallucinations or confusions it introduces? That’s similar to fuzz testing software to see what breaks or where the reasoning crashes.

Used this way, the model isn’t a source of truth or a decision-maker. It’s a stress test for your argument and your clarity. Obviously it shouldn’t be the only check you do, but it can still be a useful tool in the broader validation process.

Re: Over fifty new hallucinations in ICLR 2026 submissions

#417
post #408
post #292

Surely this is gross professional misconduct? If one of my postdocs did this they would be at risk of being fired. I would certainly never trust them again. If I let it get through, I should be at risk. As a reviewer, if I see the authors lie in this way why should I trust anything else in the paper? The only ethical move is to reject immediately. I acknowledge mistakes and so on are common but this is different leag…

Isn't this mostly a set of citation typos? To me this mostly calls for better bibtex checking, writing and checking bibtex is super annoying

Forgetting authors, misspelling them or the journals, putting a wrong digit etc... could be citation typos. I don't see how you add 5 non-existing authors and put a different—but conceptually plausible—journal in the bibtex.

Besides, I would think most people are using bibliographic managers like Zotero&co..., which will pull metadata through DOIs or such.

The errors look a lot more like what happens when you ask an LLM for some sources on xyz.

Re: Over fifty new hallucinations in ICLR 2026 submissions

#418

Earlier quoted context omitted.

Yeah seriously. Using an LLM to help find papers is fine. Then you read them. Then you use a tool like Zotero or manually add citations. I use Gemini Pro to identify useful papers that I might not yet have encountered before. But, even when asking to restrict itself to Pubmed resources, it's citations are wonky, citing three different version sources of the same paper (citations that don't say what they said they'd d…

The problem isn't whether they have more or less hallucinations. The problem is that they have them. And as long as they hallucinate, you have to deal with that. It doesn't really matter how you prompt, you can't prevent hallucinations from happening and without manual checking, eventually hallucinations will slip under the radar because the only difference between a real pattern and a hallucinated one is that one ex…

Humans also hallucinate. We have an error rate. Your argument makes little sense in absolutist terms.

Re: Over fifty new hallucinations in ICLR 2026 submissions

#419
post #292

Surely this is gross professional misconduct? If one of my postdocs did this they would be at risk of being fired. I would certainly never trust them again. If I let it get through, I should be at risk. As a reviewer, if I see the authors lie in this way why should I trust anything else in the paper? The only ethical move is to reject immediately. I acknowledge mistakes and so on are common but this is different leag…

What field are you in? In many fields it's gross professional misconduct only in theory. This sort of thing is very common and there's never any consequence. LLM-generated citations specifically are a new problem but citations of documents that don't support the claim, contradict it, have nothing to do with it or were retracted years ago have been an issue for a long time. Gwern wrote about this here: https://gwern.n…

The abuse of claims and citations is a legitimate and common problem.

However, I think hallucinated citations pose a bigger problem, because they're fundamentally a lie by commission instead of omission, misinterpretation or misrepresentation of facts.

At the same time, it may be an accidental lie, insofar authors mistakenly used LLMs as search engines, just to support a claim that's commonly known, or that they remember well but can't find the origin of.

So, unless we reduce the pressure on publication speed, and increase the pressure for quality, we'll need to introduce more robust quality checks into peer review.

Re: Over fifty new hallucinations in ICLR 2026 submissions

#420
post #363

Earlier quoted context omitted.

Talk about a buried lead... Avi Loeb is, first and foremost, a discredited crank.

That’s implied by the fact he was on the Joe Rogan show.

Please. Half of your favorite musicians, public academics, authors, industrialists, etc. have probably been on the show.
Post reply on HN