Live data from Hacker News

That post never existed. Stop listening to that thing

rachelbythebay.com

171–180 of 184 posts

Re: That post never existed. Stop listening to that thing

#171

Earlier quoted context omitted.

Can't do much about the downvote parade, but I can second this. That said, when the models are not provided the right context, and cannot fetch it for themselves, things can be rocky still. A lot less so than even just a few months ago though.

"hallucinating less" is just shifting the goal posts into vagueness. Things can be rocky still = nothing fundamentally changed. Meanwhile, if I search my file system for "foo" a trillion gazillion times, it will not return "bar" once.

If a game on release is unplayable-tier buggy, then improves over time to the point where bugs are barely noticeable, is acknowledging that going to count as "goalpost moving" to you? It never became formally verified after all, and it's even running on physical hardware... Woe are the people lying to me (nobody), the issues have not been fundamentally ruled out!

> Meanwhile, if I search my file system for "foo" a trillion gazillion times, it will not return "bar" once.

Great! Nor will an agent any more likely, cause it just sends out a tool call and surfaces its output.

It boggles the mind. One would think this is some highly secretive technology only a dozen people in the world have access to, the way one has to argue tooth and nail about trivially verifiable facts regarding it. You quite literally do not have to take either of our words or "vague" judgement for it.

Re: That post never existed. Stop listening to that thing

#172
post #119

Earlier quoted context omitted.

A “hallucination” is an authoritative counterfactual statement returned as a response. Why do you think it is impossible to engineer an LLM (by which I am including tool usage and RAG) that catches and prevents such statements?

Because they are word generators without any concept of quality save what is in their weights and they have been trained on the internet, much of which is wrong or inappropriate for any given context. They have also been trained to be people pleasers and do as they are told. The popular answer is sometimes the wrong answer.

If I cite an incorrect Wikipedia article, I didn’t hallucinate it. The citation points to a real article that happens to be incorrect.

Re: That post never existed. Stop listening to that thing

#173
post #119

Earlier quoted context omitted.

A “hallucination” is an authoritative counterfactual statement returned as a response. Why do you think it is impossible to engineer an LLM (by which I am including tool usage and RAG) that catches and prevents such statements?

Why do you think that will happen, and why do you think going for slop in the meantime could possibly bring us closer to that?

Nobody said anything about “going for slop.” You can watch today’s models actively trying to check themselves. Just use Google’s AI mode, for example. It’s far from perfect—it doesn’t fact-check every single claim, nor does it correctly understand 100% of the sources it does cite. But I’ve found it pretty useful for research as long as I use my brain and check its citations.

Re: That post never existed. Stop listening to that thing

#174
post #173

Earlier quoted context omitted.

Why do you think that will happen, and why do you think going for slop in the meantime could possibly bring us closer to that?

Nobody said anything about “going for slop.” You can watch today’s models actively trying to check themselves. Just use Google’s AI mode, for example. It’s far from perfect—it doesn’t fact-check every single claim, nor does it correctly understand 100% of the sources it does cite. But I’ve found it pretty useful for research as long as I use my brain and check its citations.

> Nobody said anything about “going for slop.”

I did. I used that phrase.

"not slop cuz pretty useful" = slop argument

Re: That post never existed. Stop listening to that thing

#175

Earlier quoted context omitted.

"hallucinating less" is just shifting the goal posts into vagueness. Things can be rocky still = nothing fundamentally changed. Meanwhile, if I search my file system for "foo" a trillion gazillion times, it will not return "bar" once.

If a game on release is unplayable-tier buggy, then improves over time to the point where bugs are barely noticeable, is acknowledging that going to count as "goalpost moving" to you? It never became formally verified after all, and it's even running on physical hardware... Woe are the people lying to me (nobody), the issues have not been fundamentally ruled out! > Meanwhile, if I search my file system for "foo" a tr…

> It boggles the mind. One would think this is some highly secretive technology only a dozen people in the world have access to, the way one has to argue tooth and nail about trivially verifiable facts regarding it.

Hahaha, I love this. The people on HN are living in some kind of bizarro world where AI is totally useless and also I guess only tried it in 2022 or something.

Re: That post never existed. Stop listening to that thing

#176

Earlier quoted context omitted.

"hallucinating less" is just shifting the goal posts into vagueness. Things can be rocky still = nothing fundamentally changed. Meanwhile, if I search my file system for "foo" a trillion gazillion times, it will not return "bar" once.

If a game on release is unplayable-tier buggy, then improves over time to the point where bugs are barely noticeable, is acknowledging that going to count as "goalpost moving" to you? It never became formally verified after all, and it's even running on physical hardware... Woe are the people lying to me (nobody), the issues have not been fundamentally ruled out! > Meanwhile, if I search my file system for "foo" a tr…

> If a game on release is unplayable-tier buggy, then improves over time to the point where bugs are barely noticeable, is acknowledging that going to count as "goalpost moving" to you?

if someone says "this hame is buggy" talking about how it might improve is moving goal posts, especially here where that "fix" is purely speculative and has not happened even once.

> Great! Nor will an agent any more likely, cause it just sends out a tool call and surfaces its output.

nope, that's so handwavy it doesn't wareant more response than that.

> trivially verifiable facts

like that game gets patches? this is too dumb for your mockery to get a rise out of me.

Re: That post never existed. Stop listening to that thing

#177

Earlier quoted context omitted.

If a game on release is unplayable-tier buggy, then improves over time to the point where bugs are barely noticeable, is acknowledging that going to count as "goalpost moving" to you? It never became formally verified after all, and it's even running on physical hardware... Woe are the people lying to me (nobody), the issues have not been fundamentally ruled out! > Meanwhile, if I search my file system for "foo" a tr…

> If a game on release is unplayable-tier buggy, then improves over time to the point where bugs are barely noticeable, is acknowledging that going to count as "goalpost moving" to you? if someone says "this hame is buggy" talking about how it might improve is moving goal posts, especially here where that "fix" is purely speculative and has not happened even once. > Great! Nor will an agent any more likely, cause it…

Continuing the gaming metaphor, what I'm trying to get at is this is like Cyberpunk 2077, and you sound like a guy who has a grand total of 0 hours in it since launch, but has developed very strong opinions about it, and refuses to accept it improved or can improve, purely because it started out so bad that that's hard for you to even imagine. Pretending to be some kind of alien, who's just going through their first exposure to practical facts somehow.

When was the last time you tried an agentic harness (Codex, Claude Code, Copilot Chat in VS Code, Cursor, Pi, OpenCode, etc.) for work in any appreciable capacity, and with what model? Surely if you're so confident they continue to be unusable and that nothing materially changed, that must be backed by a recent significant experience that way? Or even just an experience at all?

Re: That post never existed. Stop listening to that thing

#178
post #119

Earlier quoted context omitted.

A “hallucination” is an authoritative counterfactual statement returned as a response. Why do you think it is impossible to engineer an LLM (by which I am including tool usage and RAG) that catches and prevents such statements?

https://en.wikipedia.org/wiki/Entscheidungsproblem

Undecidability is not relevant here beyond e.g. a model incorrectly claiming something is decidable, and potentially even following through. It is no different to the model incorrectly claiming anything else.

These are statistical systems, so guaranteeing any particular high level behavior is not possible because of that. But given that they're working with natural language, that was never going to happen anyways, for the obvious language theoretic reasons.

The more appreciable interpretation of the claim is that they can be nevertheless tuned so that this issue becomes practically resolved. Contending guarantees and theoreticals is simply misplaced, these are not formal symbolic reasoning systems being buggy.

Re: That post never existed. Stop listening to that thing

#179

Earlier quoted context omitted.

https://en.wikipedia.org/wiki/Entscheidungsproblem

Undecidability is not relevant here beyond e.g. a model incorrectly claiming something is decidable, and potentially even following through. It is no different to the model incorrectly claiming anything else. These are statistical systems, so guaranteeing any particular high level behavior is not possible because of that. But given that they're working with natural language, that was never going to happen anyways, fo…

GP was asking for a guarantee against emitting false decidable statements. We agree that this is impossible. You can slap layers upon layers of heuristics on top, yes, but you will not eliminate all hallucinations because Church and Turing proved it fundamentally impossible nearly a century ago.

Re: That post never existed. Stop listening to that thing

#180

Earlier quoted context omitted.

Undecidability is not relevant here beyond e.g. a model incorrectly claiming something is decidable, and potentially even following through. It is no different to the model incorrectly claiming anything else. These are statistical systems, so guaranteeing any particular high level behavior is not possible because of that. But given that they're working with natural language, that was never going to happen anyways, fo…

GP was asking for a guarantee against emitting false decidable statements. We agree that this is impossible. You can slap layers upon layers of heuristics on top, yes, but you will not eliminate all hallucinations because Church and Turing proved it fundamentally impossible nearly a century ago .

I might be slow today but I'm pretty sure that the Church-Turing thesis is about undecidable statements, not false decidable statements. If a decidable statement is false... then that's it, it is just false, that's how we can decide so in the first place.
Post reply on HN