Live data from Hacker News

Over fifty new hallucinations in ICLR 2026 submissions

gptzero.me

431–440 of 442 posts

Re: Over fifty new hallucinations in ICLR 2026 submissions

#431
post #308
post #205

Earlier quoted context omitted.

IMHO what should change is we stop putting "peer reviewed" articles on a pedestal. Even if peer review is as rigorous as code reviewed (the former which is usually unpaid), we all know that reviewed code still has bugs, and a programmer would be nuts to go around saying "this code is reviewed by experts, we can assume it's bug free, right?" But there are too many people who are just assuming peer reviewed articles me…

> IMHO what should change is we stop putting "peer reviewed" articles on a pedestal. Correct. Peer review is a minimal and necessary but not sufficient step.

I agree in principle, and I think this is what's happening mostly. But IMHO the public perception of a paper being peer reviewed as somehow "more trustworthy" is also kind of... bad.

I mean, being peer reviewed is a signal of a paper's quality, but in the hands of an expert in that domain it's not a very valuable signal, because they can just read the paper themselves, and figure out whether it's legit. So instead of having "experts" try to explain a paper and commenting on whether it's peer reviewed or not, I think the better practice is to have said expert say "I read the paper and it's legit", or "I read the paper and it's nonsense".

IMHO the reason they make note of whether it's peer reviewed is because they don't know enough to make the judgement themselves. And the fallback is to trust a couple anonymous reviewers attest to the quality of a paper! If you think of it that way, using this signal to vet the quality of a publication to the lay public isn't really a good idea.

Re: Over fifty new hallucinations in ICLR 2026 submissions

#432

Earlier quoted context omitted.

this brings us to a cultural divide, westerners would see this as a personal scar, as they consider the integrity of the publishing sphere at large to be held up by the integrity of individuals i clicked on 4 of those papers, and the pattern i saw was middle-eastern, indian, and chinese names these are cultures where they think this kind of behavior is actually acceptable, they would assume it's the fault of the jour…

im not sure if you are gonna get downvoted so im sticking a limb out to cop any potential collateral damage in the name of finding out whether the common inhabitant of this forum considers the idea of low trust vs high trust societies to be inherently racist

in general, if the question is "can I divide this heterogenous population into two mutually exclusive groups based on fuzzy subjective criteria" the answer is... no

Re: Over fifty new hallucinations in ICLR 2026 submissions

#433

Earlier quoted context omitted.

Humans also hallucinate. We have an error rate. Your argument makes little sense in absolutist terms.

> Humans also hallucinate "LLM hallucinations" and hallucinations are essentially different. Human hallucinations are related to perceptual experiences not memory errors like in the case of LLMs. Humans with certain neurological conditions hallucinate. Humans with healthy brains don't. This habit of misapplying terms needs to stop. Humans are not backpropagation algorithms nor whatever random concept you read about i…

The more appropriate term is confabulate, and healthy humans do it all the time. I merely used the common, but technically incorrect term for the phenomenon in LLMs. FYI, my PhD focused on human memory.

Re: Over fifty new hallucinations in ICLR 2026 submissions

#434

Earlier quoted context omitted.

Don't understand why you're being downvoted, here.

Because the second sentence is inflammatory. The side comment is right, it's about low versus high trust societies. Even if GP made a mistake on which names are relevant, they're not being racist about it.

That's one opinion. Here's another - they were waiting with their commentary locked and loaded, and failed to even read the source material in any detail before unloading it.

They're making broad assertions about specific societies, when those assertions are in this instance in no way related to TFA.

Re: Over fifty new hallucinations in ICLR 2026 submissions

#435

Earlier quoted context omitted.

Peer review definitely does catch errors when performed by qualified individuals. I've personally flagged papers for major revisions or rejection as a result of errors in approach or misrepresentation of source material. I have peers who say they have done similar. I'm not sure why you think this isn't the case?

Poor wording on my part. I should have said "Peer review doesn't catch _all_ errors" or perhaps "Peer review doesn't eliminate errors". In other words, being "peer reviewed" is nowhere close to "error free," and if (as is often the case) the rate of errors is significantly greater than the rate at which errors are caught, peer review may not even significantly improve the quality. https://pmc.ncbi.nlm.nih.gov/article…

Thanks for clarifying, I fully agree with your take. Peer review helps, particularly where reviewers are equipped and provided the time to do the role correctly.

However, it is not alone a guarantor of quality. As someone proximate to academia its becoming obvious that many professors are beginning to throw in the towel or are sharply reducing their time verifying quality when faced with the rising tide of slop.

The window for avoiding the natural consequences of these trends feels like it is getting scarily small.

Thanks for taking the time to reply!

Re: Over fifty new hallucinations in ICLR 2026 submissions

#436

Earlier quoted context omitted.

Because the second sentence is inflammatory. The side comment is right, it's about low versus high trust societies. Even if GP made a mistake on which names are relevant, they're not being racist about it.

Yes, on looking more closely it’s possible that they made an honest mistake.

In that case, the edit button exists. It seems rather late in the day to be erring on the side of the benefit of the doubt in every case, for things like this. Much of the population is unabashedly, vociferously, aggressively racist and proud of it, these days.

Re: Over fifty new hallucinations in ICLR 2026 submissions

#437

Earlier quoted context omitted.

Yes, on looking more closely it’s possible that they made an honest mistake.

In that case, the edit button exists. It seems rather late in the day to be erring on the side of the benefit of the doubt in every case, for things like this. Much of the population is unabashedly, vociferously, aggressively racist and proud of it, these days.

> In that case, the edit button exists. It seems rather late in the day to be erring on the side of the benefit of the doubt

The edit button exists for 2 hours and this is not a person that frequently comments.

> That's one opinion. Here's another - they were waiting with their commentary locked and loaded, and failed to even read the source material in any detail before unloading it.

Well almost a day later they replied "you can google the papers and find the arxiv articles where the authors are listed". Unless that is a blatant lie, it seems like a pretty good reason to think they're using good-faith and non-racist reasoning here.

Re: Over fifty new hallucinations in ICLR 2026 submissions

#438

Earlier quoted context omitted.

Uh yeah... I would not use that tool . A tool which doesn't do its job randomly is useless.

Sorry, Utkar the manager will fire you if you don’t use his shitty calculator. If you take the time to check the output every time you’ll be fired for being too slow. Better pray the calculator doesn’t lie to you.

I’m not sure I understand the Utkar reference

Re: Over fifty new hallucinations in ICLR 2026 submissions

#439
post #93

Earlier quoted context omitted.

Indeed. The narrative that this type of issue is entirely the responsibility of the user to fix is insulting, and blame deflection 101. It's not like these are new issues. They're the same ones we've experienced since the introduction of these tools. And yet the focus has always been to throw more data and compute at the problem, and optimize for fancy benchmarks, instead of addressing these fundamental problems. Wor…

> It's not like these are new issues. Exactly, that's why not verifying the output is even less defensible now than it ever has been - especially for professional scientists who are responsible for the quality of their own work.

If I have to constantly assess every single line done by an LLM then we are fast approaching a point where it’s no longer being helpful and I’m just grading homework for a C student.

I’m not saying that isn’t what has to be done, but it kind of clashes with the whole “this will make you more productive” argument if you ask me

Re: Over fifty new hallucinations in ICLR 2026 submissions

#440
post #417
post #408

Earlier quoted context omitted.

Isn't this mostly a set of citation typos? To me this mostly calls for better bibtex checking, writing and checking bibtex is super annoying

Forgetting authors, misspelling them or the journals, putting a wrong digit etc... could be citation typos. I don't see how you add 5 non-existing authors and put a different—but conceptually plausible—journal in the bibtex. Besides, I would think most people are using bibliographic managers like Zotero&co..., which will pull metadata through DOIs or such. The errors look a lot more like what happens when you ask an…

If a person usually uses Zotero to manage literature and finds incomplete metadata when exporting BibTeX, and with the submission deadline approaching, they use GPT to complete the metadata, leading to errors, this is indeed lazy and negligent behavior. But is it what many call deceitful and unforgivable?

I believe that once this person realizes the unreliability of using GPT to complete metadata, they will no longer use such methods in the future.

I also look forward to the community's dedicated individuals developing more comprehensive automated export tools, as copying and pasting one by one is inherently tedious and should be automated.

Currently, these individuals used incorrect automated tools and placed excessive trust in them, resulting in errors. This is a profound lesson that must never be repeated.

Post reply on HN