Live data from Hacker News

The case for zero-error horizons in trustworthy LLMs

arxiv.org

11–20 of 121 posts

Re: The case for zero-error horizons in trustworthy LLMs

#13

People are going to misinterpret this and overgeneralize the claim. This does not say that AI isn't reliable for things. It provides a method for quantifying the reliability for specific tasks. You wouldn't say that a human who doesn't know how to read isn't reliable in everything, just in reading. Counting is something that even humans need to learn how to do. Toddlers also don't understand quantity. If a 2 year old…

> Counting is something that even humans need to learn how to do. Toddlers also don't understand quantity. If they're able to count to even 10 it's through memorization and not understanding.

I completely agree with you. LLMs are regurgitation machines with less intellect than a toddler, you nailed it.

AI is here!

Re: The case for zero-error horizons in trustworthy LLMs

#14

People are going to misinterpret this and overgeneralize the claim. This does not say that AI isn't reliable for things. It provides a method for quantifying the reliability for specific tasks. You wouldn't say that a human who doesn't know how to read isn't reliable in everything, just in reading. Counting is something that even humans need to learn how to do. Toddlers also don't understand quantity. If a 2 year old…

>Counting is something that even humans need to learn how to do

No human who can program, solve advanced math problems, or can talk about advanced problem domains at expert level, however, would fail to count to 5.

This is not a mere "LLMs, like humans, also need to be taught this" but points to a fundamental mismatch about how humans and LLMs learn.

(And even if they merely needed to be taught, why would their huge corpus fail to cover that "teaching", but cover way more advanced topics in math solving and other domains?)

Re: The case for zero-error horizons in trustworthy LLMs

#16

Whenveer I see these papers and try them, they always work. This paper is two months old, which in LLM years is like 10 years of progress. It would be interesting to actively track how far long each progressive model gets...

Even more interesting to track how many of those are just ad-hoc patched.

Re: The case for zero-error horizons in trustworthy LLMs

#17

> This is surprising given the excellent capabilities of GPT-5.2 The real surprise is that someone writing a paper on LLMs doesn't understand the baseline capabilities of a hallucinatory text generator (with tool use disabled).

The real suprise is people saying it's surprising when researchers and domain experts state something the former think goes against common sense/knowledge - as if they got them, and those researcers didn't already think their naive counter-argument already.

Re: The case for zero-error horizons in trustworthy LLMs

#18

Whenveer I see these papers and try them, they always work. This paper is two months old, which in LLM years is like 10 years of progress. It would be interesting to actively track how far long each progressive model gets...

Yeah well I presume at this point they have an agent download new LLM related papers as they come out and add all edge cases to their training set asap.

Is tokenization extremely efficient? Yes. Does it fundamentally break character-level understanding? Also yes. The only fix is endless memorization.

Re: The case for zero-error horizons in trustworthy LLMs

#19

Whenveer I see these papers and try them, they always work. This paper is two months old, which in LLM years is like 10 years of progress. It would be interesting to actively track how far long each progressive model gets...

Actually almost all LLMs when they write numbered sections in a markdown have the counting wrong. They miss the numbers in between and such.

So yes.

And the valuations. Trillion dollar grifter industry.

Post reply on HN